Residuals Graph¶
The residuals can be displayed in five different formats from within a TableCurve 2D Graph:
Basic Residuals - the simple difference between the Y data value and the Y predicted from the curve fit
Percent Residuals - the residuals as a % of the Y data value
Standardized Residuals - the residuals as a fraction of the fit's standard error
Distribution - the residuals in a binned histogram
Delta SNP - the residuals as a delta stabilized normal probability
These options are available in the graph toolbar and to the left side of the dialog.
By default, the Basic Residuals graph can also double as a standardized residuals graph since the Point format can specify the coloring of residuals by fit standard error.
The Residuals Graph is toggled on and off by the Residuals button in the procedure. You may also close the Residuals Graph directly. The window size and position you choose for the Residuals Graph is automatically saved across sessions.
For curve-fit residuals graphs, the runs count is displayed in brackets within the default Y-axis graph title. The runs count is the number of sign changes that occur within the residuals across the X data range. Systematic trends within residuals generally result in a lower runs count.
Residuals Distribution Graph¶
The least-squares coefficient standard errors and confidence ranges as well as the curve's confidence and prediction intervals reported by TableCurve 2D contain an implicit assumption that the residuals are normally distributed. These uncertainty statistics cannot be assumed correct unless this condition of normality is verified.
The Distribution Graph option displays a binned histogram of the residuals. Distributions with obvious asymmetry or wide tails would readily disqualify this assumption of Gaussian errors.
TableCurve 2D requires at least 16 active points in order to produce a residuals distribution. Note that any histogram is of dubious merit when data table sizes are small because of the large bin spacings. The greater the number of data points, the more accurate the distribution will be.
Delta Stabilized Normal Probability Plot¶
TableCurve 2D offers this approach as the best way to assure errors are normal. A stabilized normal probability (SNP) plot uses an arcsine transformation on both X and Y to produce a normal probability plot that uses a linear scale for both the X and Y axes. On such a plot, perfectly normal errors plot as a 45 degree line. Critical limits also have a 45 degree slope, and lay equally above and below this line.
TableCurve 2D modifies the SNP slightly and uses a delta SNP, where the X value is subtracted from the Y. This produces a horizontal y=0 for pure normal data, and horizontal critical limit lines.
TableCurve 2D plots 90, 95, 99, and 99.9% critical limit lines on the SNP plot. A 99% critical limit means that in only 1 out of 100 data sets should even a single point violate this limit. You may find the 99% critical limit the most useful. If even a single data point in the SNP violates this 99% limit, it is reasonable to assume that the errors fail this normality test. The following graph is from a curve fit which yielded normal errors.
By default, the 90% critical limit lines will be blue, the 95% green, the 99% yellow, and the 99.9% red.
You should inspect the SNP before attempting to use the parameter confidence statistics or confidence or prediction intervals in any way.
For more information on the SNP, you may refer to:
- John R. Michael, "The Stabilized Probability Plot", Biometrika, 70,1, p11-17, 1983.
- Lloyd S. Nelson, "A Stabilized Normal Probability Plotting Technique", Journal of Quality Technology, 21,3, 1989.
Maximum Likelihood¶
When the normal assumption is invalidated due to appreciable tails in the residuals distribution, equally invalidated is the assumption that least-squares is furnishing the maximum likelihood fit. In such a case, one of TableCurve 2D's robust minimizations may represent a better maximum likelihood model.
If you choose to fit a robust model, please remember that TableCurve 2D's goodness of fit statistics are all based on a least-squares common frame of reference. As such, the goodness of fit values will fail to reflect the improvement derived from switching to a robust method. Also, you should not assume simply because a distribution of errors is Gaussian that a robust procedure is if no value. Two key reasons for using a robust minimization are to deal effectively with outliers and a wide dynamic range on the y-variable.
List¶
The List Data option lists the residuals. The listing uses the TableCurve 2D text viewer facility.
Copy¶
The Copy Data to Clipboard option copies the residuals data to the clipboard. Formats include full precision binary (for spreadsheets such as Excel) and ASCII (for pasting into text editors).
Save¶
The Save Data to Disk option writes the residuals data to a supported file format. These formats include ASCII, Excel 97/2000, Excel 95, Lotus WK3, Lotus WK1, SPSS, or Systat.
MS Word/RTF Export¶
The MS Word/RTF File Export option is used to save the current residuals graph to either an MS Word file or a portable RTF (Rich Text Format) file. The graph is inserted into the file as a Windows metafile.
Analyzing Error Models¶
When the normal assumption is invalidated due to appreciable tails in the residuals distribution, equally invalidated is the assumption that least-squares is furnishing the maximum likelihood fit. If you wish to see if one of TableCurve 2D's non-linear robust minimizations would possibly represent a better maximum likelihood model, select the Error Models option. You are furnished an r² value for the four error models. This option is available only in the Curve-Fit Review.
Improving a Fit by Adding a Residuals Fit Model¶
The actual residuals can sometimes be successfully fitted. This secondary model can then be added to the primary model in order to further reduce the overall standard error of fit.
Fitting the Distribution Data¶
The residuals distribution data can be fitted by TableCurve 2D's peak functions as yet another verification of normality.
- While this approach is one further way to test normality, you are encouraged to rely primarily on the SNP plot for assuring normality.
- It is also recommended that you use this information very cautiously when there are 10 or fewer bins in the distribution histogram.
- If you choose to fit a robust model, please remember that TableCurve 2D's goodness of fit statistics are all based on a least-squares common frame of reference. As such, the goodness of fit values will fail to reflect the improvement derived from switching to a robust method.
- Do not assume simply because a distribution of errors is fitted well by a Gaussian that a robust procedure is if no value. Two key reasons for using a robust minimization are to deal effectively with outliers and a wide dynamic range on the y-variable. It is very possible that such conditions will not be reflected in the fit of the overall distribution.