Loess¶
The Loess algorithm is part of the Estimate Scattered Data option in the Non-Parametric menu. It is also used for smoothing the main data table in the Smooth Data option within the Table menu. This is the only scattered data procedure that allows for smoothing. All others offer exact interpolation.
Author
Ron Brown, AISN Software
References
This algorithm is unpublished. Information on the basic Loess concept is available in: William S. Cleveland, "Visualizing Data", 1993, Hobart Press, Summit, NJ, ISBN 0-9634884-0-6
Description
The Loess method, developed in a univariate form by Bill Cleveland is a superb smoothing algorithm, but a poor interpolatory one. A Loess interpolant contains an abundance of discontinuities. As such, this 3D Loess algorithm, developed specifically for TableCurve 3D, combines the best of the Loess type of locally-weighted fitting with the Renka I triangulation-based procedure. The Loess algorithm first performs the smoothing, and the interpolation procedure is then used with the smoothed data. This Loess procedure consists of a nearest neighbor weighted least-squares fitting to a three parameter planar, six parameter bivariate quadratic, or ten parameter bivariate cubic model.
Smoothing
The degree of smoothing is determined by the count of nearest neighbor points. A 3D-tricube weighting function is used with an integer-based cell location procedure for rapidly locating nearest neighbors. Unlike Loess, the neighbors at a boundary are not weighted to zero, but retain an influence in the fit. An order one model is planar (y = a + bx + cy), order two is a second order Taylor type polynomial (y = a + bx + cy + dx^2 + ey^2 + fxy). and order three is a third order Taylor polynomial (y = a + bx + cy + dx^2 + ey^2 + fxy + gx^3 + hy^3 +ix^2y + jxy^2). The neighbor count can vary from a lower limit that retains one degree of freedom to the total data count. The greater the data count used, the greater will be the overall smoothing. To prevent unwanted attenuation of 3D peak type data, an order 2 or 3 fit model should be used. Since the data have been smoothed, the SSE (sum of squares due to error) and r2 (coefficient of variation) are reported for the overall fit. The coloring of the data points by SE (standard error) relies on the assignment of degree of freedom based upon the local fitting function.
Interpolation and Extrapolation
The smoothed data are used as the nodes in the Renka I triangulation based interpolation procedure. This interpolant offers smooth continuous first and second partial derivatives for both interpolations and extrapolations.
Estimated Partial Derivatives:
Estimated partial derivatives are done using interpolation based upon successive gradients computed using Renka's global gradient algorithm. This approach tends to produce a very favorable accuracy for both first and second order partial derivatives. The estimated derivatives are used for graphing the partial derivative surfaces and for partial derivatives computed in the Evaluation procedure. The five partial derivatives offered by TableCurve 3D will be completely smooth for interpolation, extrapolation, and across the transition.
Estimated Volumes:
The primary method is to interpolate a uniform grid of 10,000 points and to evaluate the double integral using a bicubic B-spline integration procedure. The error reported in the Evaluation procedure will be the fractional difference with a similar 2,500 node integration. A repeat evaluation with the same limits will use an actual double integration procedure using the actual interpolant. A very fast double Gaussian quadrature procedure is first attempted to 1E-5 precision. If this is unsuccessful, a double adaptive quadrature procedure is then used.
Considerations: The Loess type of smoothing is highly effective and performance is good when used with a fast algorithm for the interpolating the smoothed data. The Renka I procedure is both fast and accurate. Although this procedure is subject to the limitations of triangulation based interpolation, a global gradient procedure helps preserve overall shape properties within surfaces.
Algorithm Adjustments
This algorithm has two user adjustments. The order determines the model to be fitted for smoothing, 1=planar, 2=bivariate quadratic, 3=bivariate cubic. The count determines the number of nearest neighbor data points used in fitting which can vary from that value which assures one degree of freedom in the fitting up to the total number of data points.
Nearest Neighbor Limitation
This algorithm is fully a nearest neighbor procedure. For the local fits to be stable, the neighbor count must be sufficient to assure at least two discrete x and y values for the planar fit, three discrete x and y values for the quadratic fit, and four discrete x and y values for the cubic fit. If the data is on a grid with a large number of observations in one variable and a small number in the other, it will be necessary to select a neighbor count that is sufficiently high to insure this stability.