50  Graphical model assessment

For this lesson, we will consider two method of graphical model assessment for repeated measurements, Q-Q plots and predictive ECDFs. We will also briefly revisit what we have done so far, the less-effective and soon-to-be-jettisoned method or plotting the CDF of the generative model parametrized by the MLE along with the ECDF. I note that we will not go over the implementation of these in Python here, but will rather cover the procedure and look at example plots. We will cover how to implement these methods in Python in the next lesson.

To have data sets in mind, we will revisit the data set from when we learned about the effects of neonicotinoid pesticides on bee sperm counts (Straub, et al., 2016). We considered the quantity of alive sperm in drone bees treated with pesticide and those that were not.

Bokeh Plot

Before seeing the data, we may think that the number of alive sperm in untreated (control) drones would be Normally distributed. That is, defining \(y\) as the number of alive sperm in millions,

\[\begin{align} y \sim \mathrm{Norm}(\mu, \sigma). \end{align}\]

We already derived that the maximum likelihood estimate for the parameters \(\mu\) and \(\sigma\) of a Normal distribution are the plug-in estimates. Calculating these using np.mean() and np.std() yields \(\mu^* = 1.87\) million and \(\sigma^* = 1.23\) million.

50.1 Overlaying the theoretical CDF

Thus far, when we needed to make a quick sanity check to see if data are distributed as we suspect, we have generated the ECDF and overlaid the theoretical CDF. I do this here for reference.

Bokeh Plot

This plot does highlight the fact that the sperm counts seem to be Normally distributed for intermediate to high sperm counts, but strongly deviate from Normality for low counts. It suggests that the bees with low sperm counts are somehow abnormal (no pun intended).

Though not useless, there are better options for graphical model assessment. At the heart of these methods is the idea that we should not compare a theoretical description of the model to the data, but rather we should compare data generated by the model to the observed data.

50.2 Q-Q plots

Q-Q plots (the “Q” stands for “quantile”) are convenient ways to graphically compare two probability distributions. The variant of a Q-Q plot we discuss here compares data generated by the model generative distribution parametrized by the MLE and the empirical distribution (defined entirely by the measured data).

There are many ways to generate Q-Q plots, and many of the descriptions out there are kind of convoluted. Here is a procedure/description I like for \(N\) total empirical measurements.

  1. Sort your measurements from lowest to highest.
  2. Draw \(N\) samples from the generative distribution and sort them. This constitutes a sorted parametric bootstrap sample.
  3. Plot your samples against the samples from the theoretical distribution.

If the plotted points fall on a straight line of slope one and intercept zero (“the diagonal”), the distributions are similar. Deviations from the diagonal highlight differences in the distributions.

I actually like to generate many many samples from the theoretical distribution and then plot the 95% confidence region of the Q-Q plot. This plot gives a feel of how plausible it is that the observed data were drawn out of the theoretical distribution. Below is the Q-Q plot for the control bee sperm data using a Normal model parametrized by \(\mu^*\) and \(\sigma^*\). In making the plot, I use a modified generative model: and negative sperm count drawn out of the Normal distribution is set to zero.

Bokeh Plot