Analog Computers

Reference / Paper · 1964

Continuous Data Analysis with Analog Computers Using Statistical and Regression Techniques

Read the PDF (10 pp) ↗

This EAI bulletin (ALAC 62023) demonstrates how fundamental statistical parameters — mean, variance, autocorrelation, cross-correlation, Fourier transform, and power spectrum — can be computed continuously from analog signals using the Exponentially Mapped Past (EMP) method and simple first-order filter circuits. The EMP approach eliminates the need to reset integrators by weighting recent data more heavily, enabling truly continuous on-line estimation. Linear and higher-order regression analysis is also developed, with a worked example applying EMP mean and variance to automated temperature control in the LD steel-making process.

Manufacturer
EAI
Year
1964
Type
Reference / Paper
Language
English
Learning track
specific applications
Pages
10
Credit
© Electronic Associates, Inc. 1964. Bulletin No. ALAC 62023.
  • EAI
  • statistical analysis
  • regression analysis
  • exponentially mapped past (EMP)
  • analog computation techniques

← Back to the Reference Library

Continuous Data Analysis with Analog Computers Using Statistical and Regression Techniques

GENERAL SECTION COMPUTING TECHNIQUES: 1. 3. 2a CONTINUOUSDATA ANALYSIS WITH ANAI,OG COMPUTERS USING STATISTICAL AND REGRESSIONTECHNIQUES ABSTRACT: This paper shows how certain fimdamental statistical parameters of a given population--the mean, the variance, the autocorrelation frmction, the cross conelation function, the Fourier transform, or the power q)ectrum--can be estimated contlnuously througtr calculations employlng simple analog techniques. tr addltlon, alalog methods are developed for contlnuous regression analysis wherein two or more populations or variables are compared statisticaUy to find significant relationships. *v v Prinr . d i n U . 5 . A . 2 6 4 O El .dr oni c As s o c i d r6 3 , I n . . 1 9 6 4 Al l R i s hr 3 R .!.N .d Bol l .i i n N o. ALAC 6 2 0 2 3 CONTINUOUS DATA ANALYSIfI WITH ANALOG COMPUTERS USING STATISTICAL AND REGRESSIONTECHNIQUES t GENERAL The need for statistical data analysis intieprocess industries is well known. Measurements ofprocess variables or parameters are subject to random disturbances such as the presence of impurities in varying amounts, environmental changes, weather, etc, Very often it becomes necessary to obtain the "best estimate" of a variableover somepriortime interral for puryoses of control. It is this concept of "estlmate" that intrcduces statistics. With the availability of small, rugged and reliable aJralog computing components especially designed for plant environments, it becomes feasible both economlcally and technically to apply statistical techniques to the analysis of continuouB data for either measurement or control. Of special importance is the fact that an alalog device can do a simple or complex calculation task while remaining a small package in terans ofphysica-I dimensions ard cost. Other computational approeches almost aIways imply the purchase ofa relativelylarge "minimum" amount of hardware. Thus, one is encouraged to e:elore the 'lsimple" applications-- situa.tions which pay their own n'ay while prcviding experience in the use and testing of the analog approach. Such computa.tions car be performed Off-Line, OnLine Open Loop, or On-Line Closed Loop on continuous-signal inputs; digitizing ofthe analog signal is unnecessary. Notsy slgnal6, unavoidablein many pilot or plant operations, whose deviation traces serse as the basis for subsequent ca.lculations can be "reduced" to more meaningful form by relative- ly simple and economica.l computer circuits. The mean and standard deviation of a noisy slgnal can be recorded continuously arrd on-line, so that all subsequentproblems in interpreting the data can be reduced markedly. More complex but still relatively inexpensive circuits can be used to record continuously, either on-line or off-line, the Fourier Series Coefficients of a signal. Or, in the dynamic testing of systems, the transformation from impulse response to frequency response can be accompllshed, thus permitting determination of the best combination of simple input and easily-lnterpreted output, THE MEAN One of the filndamental statistical estimates isthat of the t'mear" or the "arithmetic average" ofa variable or a parameter. When dealing with discrete informatlon, the mean is defined by the summation N I (1) which is recognized easily as the familiar arithmetic average. For data analysis, this statistical property is important for two reasons: 1) itis fundamental to the definitlon of other statistical parameters, and 2) lt applies equallywell to normal population distributions and to those that are not distributed normally. It would appear desirable to utilize this statistical property in the analysis of continuous data or for the measurement of continuous process variables for purposes of control. In order to do so. it becomes necessary to obtain the '.estimate', of the mean as a continuous and changing function of time. Slecifically, one must be able to define and compute the average or mean value, d, of a signal, f(t), varying with time over the interva.l Tf t_<T2. As an example, assume that a steet mill is producing a continuous metal strip which, ideally, should be uni.form but actually is fluctuating in fhickness (as shown in Figure 1) because of inevitable random disturbances in the process. Over + f(u YL Figure 2, Anolog Circuil for Colculolion oI Estimote of lhe Meon for o Fixed Time Infervol A refinement to this circuit would be to Eenerate I1t; as a continuously varying function of ti-me as shown in Figure 3. The time interval then must be considered as a variable so that the average is computed continuously from time T,. In Figure 3, T. is considered to be zero computerr time and T^ -t_z has been replaced by t since the upper limit of the integral is a variable. ul o z Yt' T ;t,r={/',t,ro, z + f 0) F LENO TH : J TI M E F igure I. Th ickne ss o f St eel Plot e f r om o Rolling M ill o s o Func t ion of Tim e any time interval of reasonable length, the mean thiclnxess should be equal to the nominal value although small deviations are allowable. A sensing device is monitoring the thiclaless as the strip emerges from the mill, and a transducer is generating a signal, f(t), which is proportional to the instantaneous thicl$ess. We would like to compute the average value of this signal so that it car be compared with the desired or nominal value in order to see if the process is ulder control. The most obvious definition of the mean or average value for f(t) over the interval T1<t<T2 is Tz-Tt f' f(tl dt t2\ '1 This value can be computed with the simple circuit of Figure 2, At time T1 the integrator is placed in the COMPUTE mode, arld at time T2 its output is observed. The integrator then call be reset and another average taken. Figure 3. Anolog Circuit lor Colculotion ofo Continuous Esfimole of the Meon for o Fixed Time lntervol Y, The circuit of Figure 3, although theoretically an improvement over that of Figure 2, has two rather obvious limitations: 1) the uncertainty of the division when t =0, and 2) the need to select maximum rtmning time in advance, since the integrator will eventually overload. This latter difficulty also occurs with the circuit of Figure 2. In each circuit, the integration can continue only over a certain range of time, and the circuit then must be reset. However, the past values of I(t) are lost in the resetting; the averag€ computed during the second !rrun" are independent of the va"luesoff(t) obtained during the first run. If a succession of runs of length T are made and the circuits are reset each time, it is clear that the last computed averag€ depends only on the behavior of f(t) in the last T units of time. !1 other words, fmm the point of view of the most recent average, information older than T units of time is obsolete. The resetting necessary with the previous two circuits ca-n be avoided, and a much simpler circuit obtained without worry about overloads and division circuits, The clue to the method is the fact that past values of I(t) become obsolete. Since the \a, basic signal, f(t), ls continuous, it seems advantageous to let past infornation become obsolete graduauy rather than abrrptly. This means that f(t) has to be defined in such a way that recent values court much more heavily tha.nearlier va.lues a^ndtlte behavior of f(t) in the remotepast has very little effect. This suggests that a weighted average be used. This The weighted-average-i(t) of a function, f(t), over T1 S t < T, with weight function C (t) is defined by ($+ wher6 d (t))0 in the interval T1<IST2. Thus The minus infinity in the lower limit serves to indicate that the average has been generated for such a long time that the effect of what happened before T1 is negligible. In otler words, since the exponential weighting function, e*', approaches zero as t* -co, the importance ofevents prior to T1 is negllgtble if T1 is suitably chosen. T^ T,' f(t) o(t) dt f( t) = (3) T^ T" c (t) dt can be simplified T^ fa i1T; = o.-.'T, | f1r1 ' =o f(t) = f: I eot f1q dt (4) -2 d,L edt T f (T) = a (7) - t) 61 r1t1 "-"(T (s) I Otterana.n (2) defines this to be the '. Exlponentiauy Mapped Past" or EMP of f(t) over a time interva.l defined by d '. Implementation of the analog circuitfor solvingthis equation is reasonably straightforward. Differentiating Equation 7 with respect to machine time, T, (t is a duIIIIny variable) gives .. T -aT = a {.(--ce ,[" .J i T^ fz l^r1 o tz e -@ .ot f1t1dt ot rlty at . ^-qT (e) t""trtotf (5) f (t) = d - - - - E- .- ..;i- o t1 -e *Numbe's in porentheses in the body o( the iext reter to rcfcrcnces I isl e d i n A P P E N D I X I, r6y at I "-oT J_"ot 1 - T Re- arranging Remembering the requirement that the recent past must be emphasized and the remote past de-emphaslzed, it follows that we should chooseaweighting function, C (t), which is increasing and such that Iim d (t) = 0' Many functions have this property but the exponential function is a natura"l one and leads to a eimple computer circuit. Pickingan exponential weighting function, eat (d>0), Equation 3 becomes (6) r1t1at Dropptng the subscripts, Equation 6 car be written as f The integral in the denominator serves to .,normalize,' the e4ression. The function O (t) can be chosen arbitrarily to emphasize or de-emphasize various parts of the interval from T1 to T2. ""t J_6 Tt o by letting Tr+-co, or d f (r) dT = "[-i(1) + r(r)] = arlq -af 1q (10) tThose tonilio. wirh lineot onalysis dnd, in porticvtoL convorurion int.sftls, wi ll rccognlze Equorion 8 os re ourpur ot o firre, whose inpulse response is aeat; thot is, o ,trsr-arder litrer wirh ti||,. Equatlon 10 is lmplemented by the simple chcult of Figure 4, which is recognizedeasily as the circuit for a simple filter or firetorderlag. Note that +t(tl a -t0t In the circult of Flgure 4 it is obvlous that an inltlal condition applied to the integrator wiU improve the computed average at the b€ginning. Thts value should represent a goodguess asto the nominal or er<pected mean value of f(t). Onenornallywould have such an estimate avallable. Ifit is a goodestimate, the computed average will be reasonable from the start; if lt is a bad one, lt wlll not mal<eany dlfference after about three to five time constants. V/. THE VARIANCE Figurc4. AnologCircuit6r Obtoining rhe EMpEstimote o{ theMeon the input and output signals have been written in terns of the more famiuar notation for time, t, which is not to be confused with the dummv variable of Equation 7. The value of the constant, d, determines how fast pest inlormEtion becomes obsolete. It is chosen arbitrarily to be large enough to filter out nonessential random fluctuations and small enough not to obscure long term trends. A useful rule of thumb can be developed by examining the response of the circuit_of Figure 4. ff f(t) chang€s abruptly (step input), f(t) wiU fouow g"aduaUy, making 95% of the change 1n 3 time constants or a time interval of 3/q.In other words, as shown in Figure 5, after ! A second important statistical parameter is the varirurce which is used to give a basic measure of the distribution of a population. It is defined as the square of the standard deviatlon and is equal to the mean-squared deviation of the variable from its mean. For discretedata, an estimateof the variance is obtained vrith the summation s 2 r. ; 12 tt-') 1 L N- i=t _-i. v (11) The term N - 1 corresponds to the number of degrees of freedom lnvolved in the calculation of the estimates of the variance (3); practical considerations dictate that the number of samples, N, will be larger than one. W FUNCTIOI{TO BE d VERAGED lrl o In a marurer similiar to the definitlon of the mean, Otterman (2) defines the EMP variance as a TEIOHTII{G FUiCTtOil c. .I I ,'g1 = .l J _@- I _---) Figure 5. The EMP Meon o{ o Conlinuous Vorioblc Provides o M.o"ut. of the Averoge o{ the Vorioblo for o Continuously Updored Fixcd Tirns lniervol. Noterhs 95% dccreose in lhe voluc of lhe weighting function over o pcriod of lcngth 3'la' This meons thot th. wlighled ovcrogc ol time, l. i5 vittuolly indepcndent of voltes thot occurrcd prior to lime t - 3/a' three time constarrts, the integrator has forgotten 95% of tbe informatlon itbad before the step change. Consequently, the EMP averag€ defined by Equation 8 is an estimate* of the mean over a time interval approximately equal to 3/a. *lt a 99% .it ''ior.ly 5/a. tion wcre used, tAe ti'.. int.Nol would be opptcxi- (r - t) dt (12) frtu - rttll' er - Y" 1t which, based on the preceding development of the EMP mean, wlll be recognized as the welghted average of the square of the deviation of the variable from lts mean, The computer circuit for calculating an estimate of the EMP variance is developed easily without recourse to mathematical manipulatlons. From the definition, and remembering that averaging is accomplished by the first-order filter circuit, the following operations are requlred: 1) form the mean with the first-order circuit. ftlter Y' 2) subtract the mean from the current value of f(t). v 3) square the difference of the mean from the current va.Iueof f(t). 4) average the square with a second filter circuit. From these requirements the circuit of Figure 6 is derived easily, r)-i(rt .,-U Figure 6. Anolog Circuit for Colculotior of rhe EMP Esrimote o{ t he Vor ionc e Example: The LD Steel Process car be used as an example of the use of the EMP mearr and variance for the control of a process. Thls is the oxygen steel making process wherein it is possible to control bath temperatures without an external fuel supply by charging the vessel with materials that are thermally balanced. The charge materials consist of hot metal (lron), scrap, and lime, The hot metal temperature can range from 2200'F to 2600'F, and, hence, it is necessary to measure the temperature of the iron to obtain a correct thermal balance. A two-color radiation pyrometer method is usedto measure the iron temperature while it is being poured into the vessel. A typical trace is shown in Figure 7. (For further details of the process the reader is referred to reference 4.) 2 600 At present, the "temperature" reading is inserted manually into the charge balance computer (the mea.n value of temperature is "guesstimated" by the operator). This could be automated easily with an EMP mear value circuit since the transducer signal is a continuous electrical signal. The condition that the reading of the mean value circuit should not be used at the treginning of tlle time history, t<t1, for reasons mentioned previously, can be automated by using the standard deviation (variance) as a control criterlon, i.e., when a21T1 rj greater than a reference va.lue, do not use f(T), when oz(T) is less than a reference value, use f(t). The reference value chosen witl depend on the maximum variance expected during smooth pour conditions. This can be mechanized readily on the analog computer by means of a comparator. with this simple technique a better estimate of the mean temperature could be inserted into the charge computer automatically ard economically. AUTOCORRELATION The autocorrelation function, defined as an integral between fixed limits, is converted easily to a continuous EMP autocorrelation frmction, O (T), by the definition T f(t) f(t -r) e-d(r' o(r) ='l J__ - t) dt 2400 2300 lO ll (13) Cross correlation a.1socould be accomplished by the substitution of a second function, g(t), into the time delay box, r, shown in Figure 8, so that the output of the delay box is g(t - i) and the output of the multiplier becomes - [rttt] [eA . ')] READINGTAKEN 2 500 these readings until the smoke has b€en blorn away (by a fan) and smooth pouring is established, Note that even after these conditions have been attained, the temperature measurement, tft<t2, is subject to fluctuations. t2 TI M E Figure 7. Plol of Hol Melol Tempcroiur€vs Time lor fhe LD Steel Process TIME DELAY [ttr][t-.r] \y The initial variations in temperature, t6(t(11, are due to the presence of smoke ard the formation of voids in the pour, The operator disregards Figure8, Circuit {or Obroiningthe ContinuousEMP Aulocorrelolion Funclion6r TirneDeloyr Reasonable time delays are obtained easily by assembling linear ana-logcomputing components. Figure 9 shows a fourth-order Paddcircuit for g€nera- F(o). The real component, E1, is formed from the transfer fuction F -L ( P +d ) d T(|\-------.,-----T .\!/ + N O T E .715r =O . 4 6 6 7 (1 ,/r) (p ra)- (18) vi c, and the imaginary component, E2, from F. "2 f(t) ol an (1e) 1p + oyz + .2 t/t Figure 9. Circuil for Fourth-OrderPode Approximotion{or ldeol Time Deloy of Mogniruder (5) ting atimedelay, r. This circuit is accurate to within 1 degree of phase shift for input frequencies in f(t) such that the product of the maximum useful signal frequency, o*, with the time delay, r, shall not exceed 6,5 radians. i.e,. < 6.5 radians 7,., -m- Note that this circuit gives P(r,r) at one value of (,. The parameter, o, can be changed merely by changing the two potentiometers labelled "@". YrJ (14) FOIJRIER AND POWER SPECTRT'M ANALYSIS The EMP Fourier transform of f(t) is defined as Figure 10. Circuit for Colc,rlolion of EMP Fouder Tronsform ond Power Spectrum T F((,) = dl f(t) e J_* or -a(T - t) -iot .. eda (15) (16) T F1<.r;= o "-jcoT I- f(t) e-a(T - t) - t) u, "jc"(T An alternative is to build similar circuits in parallel, all having the same input, f(t), and differing only in the setting of <.r. This will a1low many points of the power spectrum to be obtained simultaneouslyREGRESSIONANALYSIS The EMP power spectrum i8 defined as (17) P((') = Y,, l"t'lI ' rv--,n - t) "-o(T cos <.: (T - t dt] "'V-n" -c(T- t) si"r(T- t) dtl Figure 10 shows the analog circuit for obtainingthe power spectrurn, P(o), of the Fourier traasform, 'L Until now, only those statistical parameters that describe a single population--mean, variance, power spectrum, etc.--have been discussed. Of interest, also, is the method of statistics whereby relationships between two 01: more populations, representing different variables, are found. This method is called "regression analysis". There are many types of regression---linear, quadratic, high order, multivariate, etc. These terms refen to the type of expression used to relate the variables. For example, v y = rnx + b (Iinear regression) y = ax (quadratic regression) + bx2 + c (20) (21) \r/ Regression consists, essentially, of finding the to a set of data using a least-squares criterion. While several authors have hinted at obtaining a "least squares" fit by analogtechniques for special cases, none have shown a straightforward solution to the regression problem as defined above for continuous variables. Figure 11 shows the analog computer circuit for calculating the "least squares" parameters ,? and b using the definite integral. One should note As an illustration of how least squares fitting would be performed on the analog computer, consider the linear case defined by Equation 20, For this general equation it can be shown (6) that the following two equations will define the unlqrown par:rmeters rn and b, the slope and intercept of the line, respectively. NN lv,=-lx ,* N u Lr NN. N lxv. rt = m \/-t i=l (22) I LJ /J i=l xi1 * \- /-t i=l x-I (23) Since X and Y will be continuous fr.mctions of time. the discrete summation from l to N can be replaced by a time integral where the total time, t, is proportional to N. Therefore, Equations 22 and 23 become I. tt * =-;ft***r, (24). tt I o o r= ^[ Jo x2 d t+r/xa t (25't Jo These two equations can be solved simultaneously to yield rn and b as follows: o-It" * - ^It* * (26) v (27) 9l l -xz ra- Joxdt the"LeoslSquores" Figurell. Circuit6r Obtoining Regression Poromelers m ondb that in calculating ,? and b from the definite integral we form an r'algebraic" Ioop. This brings up the question of circuit stability. It can be shown" that once the computation is under way, (t>0), the loop will be stable unless X is constant. However, if X is constant it cannot be used as an independent variable in a correlation study. Overloads at the very beginning of the computation (a region of no interest) can be talen care ofwith feedback limiters on the division amplifiers, or by using a "steepest descent" division circuit (7). It should be observed that ,x and b are defined at every instant of tlme. For small values of t (conesponding to small sample size), the estimates of rn and b wiu be relatively insignificant and, hence, will be changing rapidly. As the time interval increases. however. the values of n a-ndb become more significant and actually should reach "steady state' or non-changing values. The technique for linear regression can be extended to quadratic or higher order regressions, it then being necessary to define a new set of equations-such as Equations 22 and 23-- fo]r determining the unknown pararneters of the regression system. Once the equations are defined, they can be converted to continuous integrals and instrumented by standard analog techniques. Again we are dealing wlth continuous signals, which means that there must be a limit of the integration interva-l if circuits such as that of FiEure 11 are to *see Reference (8J b€ used. Just as before, the need for resetting of integTators can b€ eliminated by converting the equations for rn and b to EMP equations and, thereby, obtaining truly continuous estimates of the regression parameters. Equations 22 and 23 are rewritten through N. This yields v by dividing N N Fx. Sv. L Z-/ L/ L N --- N (28) I *,', \- xi N N. i=1 i=1 N N N \- x. /Jl + h -:-:"N (2e) Figure 12. Unscoled Anolog Circuif for Colculotion o{ Continuous EMP Yolues of Regression Porcmelersm ond ! v CONCLUSIONS Recalling the correspondence between EMP variables and discrete summations, one ca.ntransform Equations 28 and 29 immediately into continuous EMP notation which gives Y= m X +b F=m X2+b X (30) (31) where m and b are no\r' the EMP estimates of the regression parameters. The analog circuit required is shown in Figure 12. The €tatements made with regard to the stability of the circuit shown in Figure 11 apply also to the algebraic loop formd in the Figure 12 circuit, It should be observed that time hae, in effect, been talen out of the problem by the conversion to EMP variables. There is no longer any need to ..reset" the integ?ators since they are now serving as convolution circuits rather than pure a.ccumulators. It follows that circuits similiar to those showncan be instnrmented for continuous higher order arrd continuous multi-variable regressions. All that is re. quired is more analog computing equipment. The conversion of statistical parameters to EMP variables enables data analysis to be performed continuously through the use of relatively simple analog circuits. Limits on the size of the integration interval norma-lly encountered with continuous signals have been eliminated; the need to ,.reset" the integrators is no longer required since they are now serving as convolution circuits rather thanpure accumulators. A continuous estimate of statistical parameters can be cafculated readily. The concept of replacing discrete summationswith the EMP mean can be a va-luableone. h addition to obvious uses for instrumentation and control, for both on-line and off-line systems, this technique a-lso can be used in analog simulation studies, For o(ample, it is sometimes deslrable to calculate the rms value of a computed variable. This is accomplished quite simply by 1) squaring the instantareous value of the variable, 2) taking the mean of the square of the variable with an EMP circuit, and 3) taking the square root of the mean. Other uses arise in simulationworkwhere Gaussian noise is used to disturb a particular parameter. Y' tl v APPENDD<I: REFERENCES (1) Davenport, W.B., Jr., arrd W.L. Root: '.A! Introduction to the Theory ol Random Signals and Noise", McGraw-HiU noot Company,ffi (2't Ottermar, Joseph: "The Properties end Mefud Mapped Past Statistlcal Valiables", IRE TRANS. on Automatic Control, Volume AC-s Number 1, January 1960, pp. 11-17. (3) Vol\ William: "Applied Stattstics fo. Engineers", Mccraw-Hill New York, 1958. p. 136. (4) Slato€ky, W.J.: t'End Point Ternperature Control Metals, Volume 12, March 19€0, pp. 226-230. Book Company, lnc., in LD Steel Making", Jouma-l of (5) Brenner, M.M., and J.D. Kermedy: "Dead Time Slmdation for Electr:onlc Analog ComNational Sirnulation Council, December:11, 1957, flE-:,, (6) Widder. D.V.: Advarced Calculus. Prentice-Hall. 194?. P. 108. (7') Favreau, R.R,! 'rDividing Circuit Princeton Comp New Jersey. Obtained by Applying qgthod of Sleepest Ascent", (8) HaDnaner, George: "Algebraic Loop8 - Some Stability CoEsideratiolEt' Educatlon and Trairilg Memo #22, Electronic Aasociate€, Itrc., PrincetoE, New Jerael'. v