Skip to content

Automated Smoothing

The Automated Smoothing option is used to smooth noisy data with minimal user intervention. This procedure is useful for rapidly comparing the six different smoothing methods offered by TableCurve 2D. The raw and smoothed data are shown in a TableCurve 2D Graph for rapid determination of the effectiveness of smoothing.

There are six methods of smoothing available:

  • FFT Filtering
  • Loess
  • Gaussian Convolution
  • Savitzky-Golay Filtering
  • Eigendecomposition Thresholding
  • Kaiser-Bessel Lowpass Filtering

Automatic Uniform X-Spacing

Only the Loess algorithm is designed to manage data with non-uniform X-spacing. The other algorithms perform best with constant X-spacing.

This automated smoothing procedure fully manages the constant X-spacing required by these algorithms. If uniformly-spaced X-values are not present, this procedure will first fit a constrained cubic spline and generate a new data stream containing the same data count. The uniform data are then processed by all of the algorithms except Loess, where it is not needed.

Optimizing Noise Removal and Derivatives

With only a single smoothing level adjustment, this automated smoothing procedure usually offers a close to optimum level of noise removal for each of these algorithms. Distortion of the underlying data trends will not usually occur until high smoothing levels are specified. For instances where an optimum noise removal is sought, TableCurve 2D offers separate Savitzky-Golay, Fourier, and Eigendecomposition procedures.

This optimization of the algorithm may be especially critical with derivatives. The derivatives available in this automated smoothing option are based on derivatives of a cubic spline interpolant of the smooth data. In general, for noisy data, derivatives based directly on the Savitzky-Golay filter are more accurate, especially beyond the first order derivative.

When derivatives are not the focus and you want the maximum in noise removal without impacting the underlying signal, consider optimizing the Eigendecomposition Denoising procedure. This can be especially effective for data sets with several hundred or more points in which a high model order can be specified. In this instance, a large number of eigenmodes can fully capture the noise component.

Generate/SAVGOL.gif The separate Savitzky-Golay Smoothing procedure offers much greater control of the Savitzky-Golay algorithm and smooth derivatives based directly on the smoothing filter.

Generate/NPARM6.gif The Savitzky-Golay Spline Estimation option additionally adds interpolation.

Generate/NPARM3.gif The Local-Regression Spline Estimation option offers a much greater control of the Loess algorithm, including higher orders, and also adds interpolation.

Generate/FOURIER.gif The Fourier Denoising option offers both frequency and signal thresholding for low pass frequency-domain filtration as well as raised data tapers to reduce spectral leakage.

Generate/FOURIER2.gif The Fourier Filtering option additionally offers channel by channel control of the Fourier domain filtering.

Generate/EIGEN.gif The Eigendecomposition Denoising option offers a greater control of the algorithm and the means to visually separate the signal and noise eigenmodes via singular value magnitudes.

Generate/EIGEN2.gif The Eigendecomposition Filtering option additionally offers the means to visualize individual eigenmodes both in the time and frequency domains. This can help determine the optimum signal-noise transition.

Level (% Smoothing)

The % smoothing for each of the algorithms defines the breadth of a smoothing or filtering window.

For the FFT Filtering, the % smoothing controls the channels that are zeroed in the frequency domain. A 10% smoothing level zeros the upper 1/2 of the frequency channels. A value of 20% zeros the upper 3/4 of the channels. A value of 50 zeros the upper 9/10 of the channels. For peak-type data, the optimum smoothing level is generally between 25 and 40%.

For the Loess and Savitzky-Golay procedures, the % smoothing determines the actual smoothing window. A 10% smoothing level results in a smoothing window containing 10% of the data points.

For the Gaussian Convolution smoothing, the % smoothing corresponds with the Gaussian FWHM being convolved with the data. A 5% smoothing levels convolves a Gaussian having a FWHM 5 times the average X spacing (sampling interval) of the data set. The Gaussian convolution also includes an intrinsic frequency domain filtration which is automatically determined.

The Eigendecomposition % smoothing specifies the eigenmode count attributed to noise and discarded. If the data count is sufficient, a model order 40 eigendecomposition is made. A smoothing level of 36% thus retains 4 eigenmodes as signal and discards 36 eigenmodes as noise.

The Kaiser-Bessel % smoothing corresponds with the normalized cutoff frequency of the filter. 0% corresponds with a frequency cutoff of 1.0, 25% with 0.1, 50% with 0.01, and 75% with 0.001.

Modify the % smoothing by entering the desired value; by changing the value with the up or down spin buttons; or holding down the right mouse button on the edit field and spin controls, selecting the level from the popup menu. After a brief delay, the smoothed data will be reflected within the graph.

AI Expert

Generate/8922.gif To make the determination of smoothing levels as automatic as possible, TableCurve 2D offers an AI Expert option which seeks to automatically determine the optimum smoothing level. This is that level of smoothing which offers the greatest possible noise reduction without adversely affecting features within the data.

Depending on the algorithm, the AI Expert is based upon a first derivative of the overall standard error (between raw and smoothed) relative to smoothing power, a first derivative of the noise reduction relative to smoothing power, or an estimation of the optimum eigendecomposition transition.

The AI Expert option will work reasonably well for many data sets. It is most easily fooled on data sets with a small number of points, and on data sets in which the noise is not consistent across the X range of the data.

Derivatives

First and second derivatives can be generated instead of zero-order smooth data. For all of the smoothing algorithms, including Savitzky-Golay, these derivatives are computed from a cubic spline interpolant of the zero-order smoothed data. While these derivatives are often very stable, the most accurate derivatives will usually come from creating a derivative filter with the separate Savitzky-Golay procedure.

FFT Filtering

This is essentially the automated version of what you can do manually in the Fourier Smoothing and Denoising and Fourier Filtering and Reconstruction options.

This option removes any linear trend which might appear as a low frequency component in the FFT, performs a forward Best Exact-N FFT, zeros the higher frequency components, performs an inverse Best Exact-N FFT, restores the linear trend, and presents the smoothed data, all in a single automated step.

Because the signal tends to appear only at low frequencies, TableCurve 2D uses the aforementioned non-linear smoothing scale.

Because of the effectiveness of the FFT, this algorithm is quite fast with large data sets.

In lowpass filtering, low frequency harmonic oscillations that are present will pass without being filtered. If you see unwanted sinusoidal behavior within the baseline areas, a time domain procedure (Savitzky-Golay or Loess) should probably be used.

Loess

The Loess procedure is a locally weighted regression smoothing algorithm that performs a full least-squares fit for each data point. Its strength rests in its ability to directly manage data with non-uniformly spaced X values. This is most numerically intense of the smoothing algorithms and very slow with large data sets.

The Loess smoothing implemented in this automated procedure fits a linear model. In general, this means there will be some attenuation of data features at higher smoothing levels.

A greater control of the Loess procedure, including higher model orders, is available in the Local Regression Estimation option.

Gaussian Convolution

This algorithm is currently unique to TableCurve 2D. It uses an automatic Best Exact-N FFT filtering for global smoothing, and convolves a narrow width Gaussian for local smoothing.

In the Gaussian Convolution smoothing, the data are transformed to the frequency domain by a forward Best Exact-N FFT, and there a Gaussian response function is convolved with the data. The frequency domain data are then automatically filtered, and the inverse Best Exact-N FFT is made.

The frequency domain filtering is automatically determined with this algorithm, and it will have some dependency on the Gaussian width used in the convolution. The % smoothing specifies only the FWHM of the Gaussian convolving the data, this as a multiple of the average X-spacing in the data set. As such, this algorithm will generally require % smoothing levels between 2 and 5, this corresponding with Gaussian response function FWHM of 2x to 5x, the average sampling interval. Anything greater is likely to produce unwanted attenuation of the data features. A response function convolution procedure does conserve overall area, however, so such attenuation is not as harmful as might be the case in a time domain procedure.

This algorithm can produce a very high degree of noise reduction, but bears the same limitations as the FFT Filtering. Since the FFT procedures are quite fast, the algorithm is efficient with large data sets, though somewhat slower than the simple FFT Filtering.

Savitzky-Golay

This time-domain method of smoothing is based on least squares quartic polynomial fitting across a moving window within the data. The method was originally designed to preserve the higher moments within spectral data. The algorithm has been modified by TableCurve 2D and will offer a higher level smoothing than that traditionally associated with Savitzky-Golay. The TableCurve 2D implementation of the Savitzky-Golay algorithm employs sequential internal smoothing passes to improve overall noise reduction.

The higher order polynomial model makes it possible to achieve a high level of smoothing without attenuation of data features. Because this algorithm relies on the linearity of an unweighted polynomial model, it is also reasonably efficient with large data sets.

Eigendecomposition

This is essentially the automated version of what you can do manually in the Eigendecomposition Denoising and Eigendecomposition Filtering options.

The smoothing occurs by performing an eigendecomposition and discarding some number of the least-ranked eigenmodes, and reconstructing the data using only the highest-ranked eigenmodes. If the data table size permits, a matrix order 40 decomposition is made. Otherwise a smaller model order is used.

Since the eigendecomposition uses SVD, it is a slower procedure than FFT filtering.

When a clear differentiation is possible between signal-bearing eigenmodes and noise-bearing ones, this algorithm offers the highest level of smoothing with the least distortion of underlying data. Unlike the other methods, it can potentially remove the entirety of random noise from a signal without touching the deterministic component.

Kaiser-Bessel

This option applies a Kaiser-Bessel time-domain lowpass digital filter to the data stream. The results are similar to the Fourier smoothing, except that a gentler transition occurs in the frequency domain.

The low pass filter's sidelobe level is set to about -60 dB and the cutoff frequency is determined by the smoothing level.

Applying a time-domain digital filter is an exceedingly fast operation, even with large data sets.

Equivalent Noise %

TableCurve 2D offers a robust noise estimation procedure that may be of some value for low-frequency signals. A cubic polynomial interpolation is made for each point using the two points to the left and the two to the right (excluding the current point). The difference between the interpolated and data values is used to generate a measure of the white noise present. This assumes that the data can be locally characterized by a smooth cubic interpolant.

The Raw Data value reports the estimated white noise in the incoming data; the Smoothed value reports the estimated white noise for the smoothed data. The percent is given as the amount of estimated noise remaining after smoothing.

This estimate does not serve as an indicator for oversmoothing. In general, oversmoothing is easily observed visually in the attenuation of data features.

The r-squared correlation coefficient is also reported. An r² of 1 is a perfect correlation while a value of 0 means the smoothed and unsmoothed data are completely uncorrelated.

List

Generate/8943.gif The List Data option lists the index, x-values, and the processed data. The listing uses the TableCurve 2D text viewer facility.

Copy

Generate/8941.gif The Copy Data to Clipboard option copies the x-values and processed data to the clipboard. Formats include full precision binary (for spreadsheets such as Excel) and ASCII (for pasting into text editors).

Save

Generate/8942.gif The Save Data to Disk option writes the x-values and processed data to a supported file format. These formats include ASCII, Excel 97/2000, Excel 95, Lotus WK3, Lotus WK1, SPSS, or Systat.

Residuals

Generate/89571.gif The Residuals button is used to graphically display the difference between the original and smoothed data.

Production Facility

Generate/8946.gif The TableCurve 2D Automation facility allows unattended processing of large numbers of data sets. The data sets can be consolidated in an Excel file or acquired using a DLL. The graphs and reports can be exported to a Microsoft Word/RTF file, while the processed data can be exported to an Excel 95 or Excel 97/2000 file.

Generate/8912.gif The Reset button restores the data to its state when first entering the procedure. If an Automation Session is in progress, the Reset button can be used to terminate the automated processing.

MS Word/RTF Export

Generate/8971.gif The MS Word/RTF File Export option is used to save the current graph to either a MS Word file or a portable RTF (Rich Text Format) file. The graph is inserted into the file as a Windows metafile.

Automatic Update

The List and Residuals windows are automatically updated when any change is made to the algorithm's settings.

Updating the Data Table

Generate/8910.gif When exiting this procedure with the OK button, an option will be presented to update TableCurve 2D's main data table with the processed data.

If Background Thread Fitting is active, the fitting will be initiated as soon as the smoothing is accepted and the data are rescanned.