Skip to content

Limitations of Least Squares

In the one-step matrix solution for the previous example, the matrix contained an x20 element even though the equation contained no more than an x10 basis function. Minimizing the square of the residuals rather than a more linear measure, such as absolute value of the difference, affects more than just computational precision.

Consider the example of two points that have exactly the same x value. Assume each point also has the same weight or standard deviation, meaning each is weighted equally in computing the sum of squared residuals. If one of these points is twice the y-distance from the curve as the other, its influence in this sum of squared residuals will be four times greater than that of the closest point.

Outliers

This raises two important limitations to be aware of in least squares fitting. The first is that outliers can significantly influence the fitted surface and its coefficients.

By Standard Error

TableCurve 2D offers a variety of ways to recognize such outliers. When points are colored by their relationship to the overall standard error of fit, it is easy to spot points that lay outside 3 standard errors. In the default colors, such points will be red. Such points can be turned off and the data refitted without them. You may also wish to consider the merit of disabling points that lay outside 2 standard errors of the fit. These points have a yellow color by default and represent points which lay outside 95.4% of the error distribution for the fit.

You must keep in mind that this coloring by standard error is based on the overall least squares solution. This does not take into account the magnitude of the y-value nor how accurately a curve may be determined in any particularly region.

For example, consider a transition curve that begins at y=0 for the initial x values and which rises to y=100 for the final x values. Let us assume the standard error for this curve fit is 3.0 and that the curve is very accurately determined at the lower and upper plateaus, but very poorly in what is a very sharp transition.

All points will be colored based upon this 3.0 standard deviation for the fitted line. The accurately determined points with a y value near zero will be treated in exactly the same way as those accurately determined at the y=100 upper plateau and these will be treated exactly as those marginally determined points with a y-value near 50 in the middle of the transition. If the observed y value for a point is outside this ±9, it is colored red and said to be outside 3 standard deviations for the fit. If it is between ±6 and ±9, it is colored yellow, regardless of the magnitude of the y-value, nor how accurately the curve is determined at the x of concern.

Perhaps a point should genuinely be an outlier on the lower or upper plateau if its y-value is ±1 of the computed value. Such points do not appear as outliers because of the large impact of the points inaccurately determined in the transition.

By Confidence Intervals

The confidence intervals for a curve measure, at a specific x, the y range associated with a specified probability for the true y value. It is used most often when data consists of averages and inverse variance weights, or when data consists of multiple appended sets. Prediction intervals for a curve are likewise measured at a given x, and consist of the y range associated with a specified probability for the y value of the next experiment. Prediction intervals are commonly used when data consists of a single experiment, and you are interested in predicting the values for the next experiment.

Prediction intervals are also a useful method for discerning and disabling outliers in the data. Since an interval is a more local measure of confidence, whether or not a point lay outside a given probability also depends on how effectively the curve may be characterized by the curve-fit in that data region. As a rule, a curve-fit tends to be best determined near the mean of the x values and less so near the edges of the data region.

Whether outliers are removed by standard error or by prediction interval seems largely a matter of taste. In most cases, points outside a 95% prediction interval will also lay outside two standard errors of fit.

By Error Bars

If the data has been constructed from replicate measurements, either by external compilation or from within TableCurve 2D, errors bars will be available. These can be used to spot individual measurements which may be in question and should not be included in the composite y values.

Y Data Spanning Orders of Magnitude

The other major limitation in least squares occurs when the y-data spans many orders of magnitude. Imagine a data set whose y axis is on a log scale ranging from 1E-5 to 1E+5. An error of 1E-6 may be typical at 1E-5 values whereas an error of 1E+4 might be typical at 1E+5. The squared residual for the typical error at y=1E-5 would be 1E-12 whereas the squared residual for a typical error at y=1E+5 would be 1E+8. When 1E+8 is added to 1E-12, even with Intel 80 bit hardware floating point, the result is still 1E+8. The error at the low y-value is insignificant in the overall sum of squared residuals. As such, the curve fit is unlikely to fit the data at these low y-values.

Weighting for Magnitude of Errors

When errors vary as a constant fraction of the y-value, one solution is to weight the data for the magnitude of y. In TableCurve 2D, the simple calculation W=1/(Y*0.05)^2 would adjust for y magnitude when the errors are approximately 5% of the y magnitude. This type of weighting scheme is rather dependent on having a constant fractional error.

Another solution is to simply plot and fit data that has been pre-transformed to ln(y). In TableCurve 2D, all statistics and all fitting are based upon y. Even when an equation is y-transformed, the goodness of fit statistics are computed based upon y rather than upon the scale of the transform. TableCurve 2D must assume that the y variable represents the world scale of the data. For equations to be comparable, this means that all data must be based upon a common y-measure of error. One equation cannot report statistics based on minimizing the ln(y), another or minimizing the inverse of y, and yet another on minimizing y itself. All measures of a curve fit must be scaled to the dimensions of the original y-variable.

Log(Y) or Ln(Y) Pre-Transform

This is where you can help yourself free of the limitations of the least squares process. A simple Y=LOG(Y) calculation will convert the data from a Y scale to a base 10 logarithmic scale. The 1E-5 to 1E+5 data that spanned ten orders of magnitude now spans from -5 to +5. Weighting suddenly becomes less critical since the errors at all data points will significantly influence the overall least squares minimization.

The problem is solved by changing the scale of the problem from one that plays havoc with least squares and the limitations of floating point hardware to one that utilizes nearly all of the precision in the machine.

1/Y Inverse Transform

Often times an inverse transformation will work quite effectively in changing the scale of a problem. For data beginning at y=1 and increasing to very large y values, a calculation Y=1/Y will result in all Y data being between 0 and 1. This is the optimum dynamic range for floating point processing.

As a general rule, if your data spans only a few orders of magnitude, you should find TableCurve 2D’s y-based least squares minimization to be fully satisfactory. Once your data starts to span more orders of magnitude, you should seriously consider assisting TableCurve 2D by redefining the scale of your problem. You may also find that by minimizing the squared residuals in ln(Y) or 1/Y rather than Y itself, you find a host of additional linear equations that you would not have otherwise realized.

For non-linear equations, and for those linear equations you might be willing to fit via a UDF, there is another especially attractive alternative. By using the robust minimizations available in TableCurve 2D’s non-linear fitting engine, it is possible to deal with outliers and Y-dynamic range in a very effective and automatic fashion. It should be noted that non-linear robust fitting represents a very non-traditional approach to curve-fitting because least squares is altogether abandoned in favor of minimizations which are far less sensitive to outliers and Y-dynamic range.