SIMSTAR — An Attached Multiprocessor for Dynamic System Engineering
Bulletin No. 04
July 1983
SIMSTAR — AN ATTACHED MULTIPROCESSOR FOR DYNAMIC SYSTEM ENGINEERING
PRINTED
This paper was presented July 13 at the 1983 Summer Computer Simulation Conference in
Vancouver, British Columbia, Canada.
| Electronic Associates, Inc.
185 Monmouth Parkway, West Long Branch, NJ 07764 (201) 229-1100
IN. U.S.A.
EA| Electronic Associates, Inc.
SIMSTAR - AN ATTACHED MULTIFROCESSOR
FOR DYNAMIC SYSTEM ENGINEERING
Hy: J. Paul Landauer
Flectronic Associates,
West Long Branch, NJ
ABSTRACT
ee eee ee ee ee Gee ee oe
This paper describes a new multi-
processor that has been developed by
Electronic Associates, Inc. for scientific
analysis of dynamic systems. Proven par-
allel and sequential computing methods are
integrated im SIMSTAR(TM) to provide a
unique capability for mixed continudus/
/Gdiscrete system simulation and signal
Processing. In contrast to the earlier
manually programmed analog and hybrid
computers, SIMSTAR is a completely auto-
matic device driven fram high level |
software in a Host data processing
computer.
| The requirements for dynamic system
simulation in different fields is developed
in the paper to establish a criteria for
measurement of SIMSTAR performance relative
to alternate computers. New concepts in
system architecture, component technology,
and system communication features are |
described. An overview of the SIMSTAR
programming system is covered to see the
flow from the simubkation language input’ to
the program segments for the various
processors. Also, a brief discussion of
program operation from the Host terminals
through a run-time executive is described.
| SIMSTAR is a new computing tool for
engineering analysis of dynamic systems.
By combining the latest linear and discrete
integrated circuit technologies, the
earlier hybrid computer concepts have been
extended into an automatic, high perform
ance device which can be attached to a
range of medium scale data processing
systems. Initially, the Host system is one
of the GOULD S.E.L. 32 series. Applications
OF SIMSTAR include conception, evaluation,
and optimization of dynamic physical
systems in ali of the engineering fields.
By integrating a high-speed, economical
digital arithmetic processor with a unique
automatic, stored-program parallel
processor, equivalent computing speeds over
200 million operations per second are
obtained. |
Inc.
O77 64
The parallel processing system can be
expanded to over 400 mathematical computing
blocks which are interconnected by a solid-
state switch matrix. Functions implemented
in these blocks include continuous integ~
ration, linear arithmetic, non-linear
operations, and logical processing. For
most applications, the digital arithmetic
processor is assigned those equations which
represent the slowly changing environment.
The system, being designed to operate in
this environment, is usually modeled on the
parallel processor which can automatically
handle state variable discontinuties,
simultaneous algebraic relationships, and
natural frequencies over 1 kHz. The
complexities of sophisticated numerical
integration techniques can be avoided by
use of this combined approach.
Since the SIMSTAR Multiprocessor is
completely automatic in aperation, it can
be programmed in the same fashion as an
automatic data processing machine. The
user may prepare programs either in FORTRAN
or in a high level Continuous System
Simulation Language. Setup and interaction
with SIMSTAR is handled by the built-in
digital arithmetic processor.
The design and evaluation of complex
dynamic systems has been one of the most
challenging engineering tasks since the
time of Sir Isaac Newton. Many different
mathematical and computer means have been
developed as aids. The primary math-
ematical tool has been modeling by systems
of ordinary and partial differential equa
tions along with supplementary algebraic
equations. When these models are solved,
the process is often called "Simulation, ”
although this term is rather ambiguous,
since it is also used for simple animation
of an environment for display or training
ourposes.
During the last twenty-five years, the
computer tools for solving these models
have been primarily combinations of
electronic analog and digital machines
which are usually termed hybrid computers.
Farly analog computers used mechanical
devices which were relatively unreliable.
The first digital computers were very slow,
expensive, and relatively difficult to
program. Though the 1970*s, both of these
technologies evolved so that powerful,
economical digital, analog, and hybrid
systems became available. During this same
span, the complexity of applications
increased at the same rate so that today
simulation is still the most demanding
computer application.
SIMSTAR is a quantum step forward in
computer technology to meet the increasing
demands of engineering simulation. It can
be effectively used in large-scale sim-
ulation laboratories in the fAlerospace,
Nuclear, and Electrical fields as well as
central scientific/engineering computer
facilities in high technology companies.
Convenience features available on the Host
computer such as color graphics, network
access, and Computer Aided Engineering /
Design tools can be used with the SIMSTAR
attached processor. Accordingly, complete
integrated system desiqn studies can be
carried-out amongst different divisions of
campanies. |
GENERATORS CONTROLLERS FILTERS SCA'S
ELECTRICAL SYSTEMS
TRANSLATION ROTATION SURFACES ACTUATORS SEEKERS
MISSILES
¥
o
we
m4 +
3
2 3 SHORT PERIOD CONTROL
< PHUGOIO PITCH STRUCTURAL SURFACES
ert + t + —
us & AIRCRAFT
w =
nu
an | + + t + —_—_——-——|
Ps Z 0.03 0.3 3 0 300 3000
a am SUBSYSTEM NATURAL FREQUENCIES ~ HEATZ
&
8
b {}——_-—}
0.01 0.1 1 10 100 1906
SPEED REQUIRED — MILLIONS OF NGRMALIZED
OPERATIONS PER SECOND
Figure 1 - fhe Application Requirement
chet cinteke Se aie 2 ee ee eee eke deity Gee ie GE ees ees Ge wake Ge Gps ee ee eee ee eee
A display of the requirements in three
major fields of application is shown in
Figure 1. The abscissa of this chart is
divided into five decades of computer speed
in terms of millions of "Normalized
Operations-Per-Second" (NOPS), which is a
simple method of comparing processor per-
formance for this class of application. A
Normalized Operation (NOF) is essentially a
Simple memory reference instruction such as
LOAD or ADD. Multiply or Divide are usual-
ly counted as 3 NOFs. If we assume that we
need to solve a system of 25 ordinary
differential equations with appropriate
non-linearities to three place solution
accuracy, the equivalent natural frequen
cies of the system can be determined as
shown in the chart for the various speeds
shown. That is, real-time simulation of
this system operating at a natural frequen—
cy of about 3 hertz will require a digital
processor performing approximately one
miliion NOFS. If the frequency increases
to 30 hertz, the equivalent processing
speed will go ta 10 million NOFS for the
same problem complexity and accuracy.
Assuming a single integration method, the
speed requirement is proportional to the
natural frequencies of the subsystems being
modeled.
For aircraft, the Phugoid Mode cal-
culations operate at about 9.01 hertz. The
Ghort Feriod Pitch frequency is about 1
hertz, Structural dynamics are in the range
of 20 hertz, and the Control Surface
deflections will range from about 50 to 100
hertz. If the Control Surfaces are rep-
resented by a system of 25 transfer func-
tions with limiters, the real-time solution
will require about 20 million NOPS. For
missile system simulations, the frequencies
also vary thoughout this entire range from
the Translational equations at less than
O.? hertz to the small infrared or radio
frequency Seekers at over S00 hertz. In
between are the Rotational equations at
about 2 hertz, the Control Surface dynamics
at about 20 hertz, and the Actuator
response which is about SO hertz.
Another important application for
SIMSTAR is modeling electrical power sys-
tems for power inverters, battery storage
systems, and general power control of
rotating machinery. Again, the frequencies
vary from the mechanical time constants of
generators through the speed controllers,
the line filters, and the Silicon
Controlled Rectifiers (SCR) which are in 7
the DC-AC inverters. Froper representation
of these systems can require a computing
capability of over 100 million NOFS.
EXTERNAL FACILITIES
(OP TIONAL} FUNCTION 8 {
7~T TT aeae GENERATION .
- |
PROCESSOR 14
1
_
— se eer ew eee ee
PARALLEL
SIMULATION
PROCESSOP.
PARALLEL DIGITAL
SIMULATION ARITHMETIC
PROCESSOR PROCESSOR
a
PLU
(OPTIONAL
EXPANSION}
MEMORY PORT
HOST
DATA
PROCESSING
COMPUTER
LISTINGS
USER TERMINALS
Figure 2 — The SIMSTAR System Architecture
MULTIPROCESSOR ARCHITECTURE
A Block diagram of the SIMSTAR dual
muiti- processor attached to a Host
computer 1s shown in Figure 2. In most
real-time simulation laboratories, the
SIMSTAR model must interface to various
external facilities to test actual sub-
systems or for human interaction. Various
types of continuous display and recording
devices may also be used. Parallel analog
and binary signals as well as a digital
data port are available.
The basic SIMSTAR multi-processor,
which is attached to a Host computer, is
composed of a single Farallel Simulation
Frocessor (PSF) and the Digital Arithmetic
Frocessor (DAF). The second Parallel
Simulation Frocessor and the Function
Generation Frocessor are optional devices
in the system and provide parallel
extension of the computational power.
Included in the FSP is a Parallel Logic
Unit (FLU) and a Parallel Math Unit (PMU),
which provide the heart of the computing
POwer of SIMSTAR. Sequential digital
computing is provided by the DAP composed
Of a 32 bit CPU and MOS memory. With an
Optional Floating Point Accelerator (FPA),
this processor is approximately equivalent
to a VAX 11/780. A basic setup and mon-
itoring interface capability between the
FSP and the DAP is provided as an inherent
part of the minimum SIMSTAR system. In
addition, a Data Conversion Frocessor (DCF)
can be added to the DAF for high
speed/accuracy direct memory data
communication with the PSP. With this
interface, the PSF can also communicate
data directly to/from the Host at high
speed if it is needed as a part of a
computational task. All programming of the
DCF is done using Channel Programs setup by
the DAF.
The user operates from terminals on the
Host and uses the normal file handling
Capability of that system. Also, high-
level software to load the SIMSTAR
processor executes on the Host, and
apprapriate listings can be obtained. The
object program can then be downloaded into
the DAP and the FSF to operate the sim-
lation. The DAF, in turn, will load the
Function Generation Processor (FGF) with
appropriate data and programs.
In most applications, one of the PLU’s
will act as the master timing control for
the entire multiprocessor program, since it
has a programmable crystal clock and
Interval timing capability built-in. This
device interrupts the DAF for time-critical
processing. Data produced is stored in the
DAF memory from which it can be accessed by
the Host for display on graphics terminals,
listings, graphical hard copy, or permanent
* storage on the disk.
The subsystem models developed in the
research and development department on
SIMSTAR can be made available to mary
design engineers by establishing an appii-~
cations library on the Host digital data
processing system. These Libraries may be
further used by an Executive routine to
provide Froblem Oriented Languages. In
this way, design engineers can make
effective use of Simulation methodologies
without the need to develop new math~-
ematical models each time.
SIMSTAR r
+ FPA
DAP HOST
= te : +
¥P-3300
an |
MULTIPLE
SIMSTARS
SPECIALIZED
DIGITAL
1 10 100
Lh —t t=
0.01 0.1
AP-120B| AO-10
DIGITAL MC 68000 e088 12 VAM A 32 CRAY CRAY
COMPUTERS
8087 37 «614/780 8) 6780 i"
SPEED — K WHETSTONES
1000
t 10 100
SPEED — MILLIONS OF NOPS
Figure 3 — Comparison to Typical Commercial
Computers and Processors
a ee ee ee ee ee ee eplcke cities Ah ich es oaks eee eee eee eee wees Ee eee OO ee ee Gee cee ee ee eee
The next chart, shown in Figure 3,
presents typical commercial computers and
specialized digital processors as compared
to the SIMSTAR attached multiprocessor.
The abscissa is again the speed required in
Normalized Operations—-Per-Second ranging
from 10,000 to 1 billion. The corres-~
ponding speed in Kilo Whetstones is also
shown om this chart since most computer
performance is quoted in terms of the
Whetstone benchmarks. A speed of i million
NOFS is about 700,000 Whetstones.
The various computers being considered
throughout this performance range go from
the Motorola 68000, which is a 16/32 bit
LSI processor, to the Cray II, which is a
quadruple 60 bit floating-point machine
with 4 nanosecond cycle time. Costs of
these digital processors range from about
$4,000 for the 68000 to about $25 million
for the Cray II. It can be seen that the
DAF covers the range from 10,000 NOFS to
about 1 million NOPS when it is enhanced By
the Floating-Point Accelerator (FFA). The
equivalent speed of the DAF with the FFA is
about 670k Whetstones. In contrast, a VAX
11/780 1s about BOOK Whetstones. If
digital processing performance greater than
this is required, the user can apply the
Host digital computer to the simulation.
The chart shows that one could use a GOULD
S.E.L. 32/8780 dual processor for this
function to increase the digital floating-
point performance of SIMSTAR to about 8&8
million NOFS.
The Farallel Simulation Processor 1s
the only practical device for speed
requirements from 20 million NOPS to 200
million NOPS. Of course, the PSP can also
be used at lower speeds to overlap with the
Host or the Digital Arithmetic Processor.
If speeds greater than 200 million NOFPS are
required, one can perform these caicula-
tions on the FSP at somewhat reduced
accuracy without concern about numerical
instability. Applications requiring up to
nearly 10 kHz can be handled in real time
on the PSF if one can live with up to SZ
error. In the future, it 18 planned that
the system could be expanded to 3 SIMSTAR
dual multiprocessors ina system for a
combined equivalent performance of over 50a
Million Normalized Operations—-Fer-Second.
Obviously, one would not use Cray
computers in this type of application since
the cost 1s prohibitive. However, a possi-
ble alternate is one of the specialized
high-speed digital processors. For example,
the Applied Dynamics AD-10 can be viewed to
Operate in the range of 10-20 million NOFS.
The sixteen bit fixed point operation may
not be sufficient since one often requires
greater resolution and range for slower
Variables. A new floating-point version of
the AD-10 has been announced, but no
systems are yet installed. The digital
arithmetic processor in SIMSTAR provides
real-time 32 bit floating-point capability
for variables and 64 bit precision for
integration.
Other high-speed specialized digital
processors which have been used to perform
various simulation tasks include the
Floating-Point Systems AP-120-B and the
FPS-100. These devices are somewhat slower
than the AD-10 but have full floating-point
computation capability and some specialized
software useful for simulation.
The VP-3300 is a version of the
earlier MAP-300 manufactured by CSPI.
unit operates through a Common Memory
interface to GOULD S.E.L. 32 computers
which eliminates the usual program and data
transfer overhead. Specialized software
for scientific computation is also avail-
able for this device. A version of this
processor can be added to a SIMSTAR system
for function generation or coordinate
transformation.
This
All of these digital devices can
assist in real~time simulation only to 20
or 30 hertz. Above this, the SIMSTAR
Parallel Simulation Processor is the only
practical choice. Of course, if one can
justify not modeling the higher frequency
terms, the slower devices can be used.
It can be seen in Figure 3 that
SIMSTAR is designed to provide the
appropriate type computing devices for the
accttracy and speed requirements needed in
many large scale simulations. If a partic—
ular model has lower frequencies and does
not demand real time, the problem can be
time-scaled upward to take full advantage
of the extremely high-speed performance of
the PSP. This excess of computing power
simplifies the task of the simulation
engineer since he does not have to optimize
the utilization of the processors or be
concerned about complex numerical
integration methods.
TO HOST DIGITAL
COMPUTER
MEMORY PORT
TODAP
MEMORY SUS
. _finwmane |
SIMBUS
| PROCESSOR
INTERFACE
SIMBUS
PROCESSOR
| INTERFACE
LOCAL
CONTROL
PROCESSOR
SIMBUS . 16 BIT CATASCONTAOL BUS
| PARALLEL LOGIC UNIT
"COMPARATORS [>
a SIGNAL
PROGRAMMABLE foseececneracean
4 MATRIX
LOGIC
, CLOCK 7
ARAAYS
SETUP LOAD
PARALLEL MATH UNIT
MATHEMATICAL
COMPUTING
BLOCKS
BLOCK
CONNECTION
MATRIX
Figure 4 —- Functional Block Diagram of the
SIMSTAR Parallel Simulation Frocessor
PARALLEL SIMULATION PROCESSOR
The Parallel Simulation Processor
(PSP) is the main computing device in the
SIMSTAR system. It incorporates high
speed, continuous Mathematical Computing
Blocks and a set of Programmable Logic
Arrays to solve a model composed of non-
linear differential equations combined with
Switching and control as needed. The
concept of dedicating computing devices to
specific terms in particular equations has
been successfully used for many years in
analog and hybrid computing methodologies.
SIMSTARK incorporates these proven concepts
into ‘a totally new machine which provides
automatic operation from a Host digital
computer.
A functional block diagram of the FSP
as implemented in the SIMSTAR system is
shown in Figure 4. The Farallel Logic Unit
1s composed of the Logic Siqnal Matrix and
the Programmable Logic Arrays. Mathemat—
ical Computing Blocks combined with the
Block Connection Matrix make up the
Parallel Mathematical Unit of the PSF.
double lines represent parallel commun-—
1cation paths to handle many mathematical /
logical variables simultaneously. Selected
cantinuous mathematical variables produced
by the PMU can go to the Comparators which
test for a specified threshold level. If
the variable exceeds that level, a logic
signal is generated through the Logic
Signal Matrix to implement control func—
tions in the PLU. Finally, the logical
states out of the FLU drive the switching
devices and mode sequencing of the Math-
ematical Computing Blocks.
The
Interprocessor communication between
the Local Control Processor (LCF) and the
various computing and monitoring systems in
the PSP is handled by the 146 bit SIMBus.
This iss an adaptation of the Intel Multi-
bus. The SIMBus Processor Interface (SFT)
units provide memory mapped communication
between the Host, the DAP, the RAM an the
SIMBus. and the various storage devices in
the PMU and PLU.
Loading and setup of all components in
the PSF is handled by the LCF. That is,
conversion from floating-point data in the
Host or the DAP to the optimum gain set-
tings and configuration selection 15 per-—
formed automatically by the LCF. Unused
elements of each hardware macro are |
defaulted out so the user need not be
concerned. System monitoring for over—
range and siqnal instability is handled by
the LCF. Also, the autorange intelligence
in the ADC to provide high resolution
readout is performed by the LCF. Finally,
the Local Control Processor performs the
periodic automatic device calibration in
the PMU.
Parallel Mathematical Unit
In addition to automatic, high-speed
computing capability, SIMSTAR includes
innovative component and system design to
ensure accurate results over a wide dynamic
range. Essentially, the same concepts used
for floating-point digital computation
using REAL numbers is employed inside the
computing blocks of the Farallel Math-
ematical Unit. A continuous computing
device such as the PMU employs electrical
signals to represent dynamically changing
variables. All information may be con-
sidered normalized so that the maximum
value of a signal is near unity. Although
an exponent cannot be carried with the
signal between math blocks, within a block,
pseudo floating-point methods are used to
provide full four digit accuracy over a
greater range than ever possible before.
In particular, non-linear blocks which
produce the product, quotient, or square
root of signals employ an automatic change
of exponent to extend the accurate dynamic
range by 16-to-1. This operation 16
completely transparent to the user.
Since a signal is accurate to about
five digits, this provides much greater
range than used in the earlier analog
computers. The monitoring system uses an.
Autoranging feature to allow accurate
readout of smaller signal values. PFara-
meter data loaded into coefficient devices
is in floating-point and the built-in
microprocessor firmware establishes an
optimum combination of binary fraction
value and gain for eachorun. jategrator
gain can range from 10 ~ to 10 ~ with full
four digit resolution. That is, the user’s
program simply sets the gain as a floating—
point number in this range and the appro-
priate internal settings, gains, and
integrating devices are selected by the LCF
for optimum accuracy.
All of these pseudo floating-point
features combine to make the SIMSTAR system
much more User Friendly than any other
device for simulation. Furthermore, the
accuracy of solution for typical appli-
cations is at least ten times that
available with the earlier hybrid
computers.
I
Parallel Logic Unit
Froblem timing and control is provided
by the PLU in combination with the built-in
crystal clock and the timing registers
shown in the feedback around the PLU. Of
course, a continuous variable representing
time in the problem can be produced using
one of the Integrating computing blocks.
The implementation of sequential logic in
the PSF is done by using the delay flip-
flops (F/F), which are provided in the
feedback of the PLU as shown in Figure 4.
The SIMBus/Local Control Processor
Setup of the PSP and digital run-time
communication between processors is per-
formed through the SIMBus, which is a 16
bit high-speed implementation of the Intel
Multibus. This bus is also used to commun-
icate with the Host digital computing
system in a memory mapped fashion. The
tocal Control Frocessor (LCP) with its
associated Firmware is shown above the
SIMBus in Figure 4. This device translates
commands from the user program referencing
Math Computing Blocks to the actual
implementation of those functions on the
hardware macros of the PSF. The LCP pro-
vides the on-board intelligence in the PSF
for any binary communication required from
user programs to the physical computing
subsystems. Alsa, a Random Access Memory
(RAM) on the SIMBus is loaded by the LCP
during the setup or operational phases of
the PSP. This memory, as well as many
registers throughout the PSP, can be
accessed through the SIMBus Frocessor
Interface (SPI) which is a 32 bit memory
port to the Digital Arithmetic Processor.
As shown in Figure 4, a second Simulation
Processor Interface is used to connect the
SIMBus to the Host digital computer memory
bus. This allows the Host to read the
binary status of the entire PSP and store
it on disk. To restore a program, a
transfer from the disk file back into these
memory locations will re-establish the FSP
operation very rapidly.
For readout of the signals from the
Mathematical Computing Blocks, an auto-
ranging Analog-to-Digital Converter (ADC)
is attached to the Block Connection Matrix.
The LCF can acquire this data and translate
it into appropriate formats for floating-
point transfer back to either the DAP or
the Host.
PSP Computing Performance
As discussed under Applications,
engineering simulation studies demand
computing performance which exceeds that
available with any standard digital
computer system. In SIMSTAR, this speed
requirement is met by employing a large set
of parallel computing blocks which operate
continuously upon signals. Earlier in this
paper, the equivalent performance of a
SIMSTAR attached multiprocessor was quoted
up to 200 million Normalized Operations—
Per-Second (NOPS). This was based upon a
typical mix of components on a unit having
dual parallel processors operating at an
average solution frequency of 300 hertz.
DIGITAL
OPERATION
HAROWARE
DO/SUB/AND
(
2S, MULT/DIV
ret ty depefe ~
a NOP’S PER
cp PASS
MACRO
< .
LIMITED 7 44
INTEGRATOR/TS 2 :
| SUMMER/LIMIT 2 [3 |
MULTIPLIER
(DIV/SQRT)
3 INPUT MULT | - 2 |
n ef
@) we fe fe
SPAT a
NOPS - Normalized Operations per Second
Figure S&S - Equivalent Digital Operations
for SIMSTAR Hardware Macro Components
This equivalent performance 18 derived
(1) by estimating the number of normalized
operations required for a digital processor
to perform the same computations as each of
the components of the SIMSTAR parallel pro-
cessor. The chart in Figure S shows nine
basic component types including functions
of two (or more) variables with the NOFs
required for a single integration pass and
a fourth order Runge-kutta integration
step. If a single pass predictor/
corrector integration method were adequate,
the NOPs per Fass could be used. However,
this type of method can introduce large
errars for discontinuities in state
Variables.
COMPONENT TYPICAL EQUIV. TOTAL
TYPE #/PP NOP/RK4 NOP/RK4
LIMITED
INFEGRATOR/TS
SUMMER SWITCH
SUMMER LIMIT 1440
576
MULTIPLIER
DIV/SQRT
THREE INPUT
MULTIPLIER
SINE/COSINE
f(2} or £(3), £(4)
79
Le
1824
ad
eS
S
COMPARATOR
PARALLEL LOGIC
UNIT
15372
TOTAL NOP/RK4.
AT 20 STEPS/CYCLE (0.1% ERROR) 300,000 NOP/CYCLE _
AT 300 HERTZ (AVG.) (0.1% SOL ERROR) 90 MILLION NOP/SECOND)
Figure & —- Single SIMSTAR Parallel
Frocessor Performance
NOP’S PER
& | ster
«onl
On the next chart (Figure 6), the
total SIMSTAR PSP performance is calculated
based upon the typical number of each of
these components which can be programmed in
parallel. Including 500 NOFPs per RK-4 step
for the PLU, over 15,900 normalized oper-
ations are required per integration step.
If we assume 20 steps/cycle, RK-4 integra-
tion will result in about 0.1% error. A
digital processor would need to compute
390,000 NOFs per solution cycle to match
the performance of a fully expanded PSP.
if we assume an average frequency of 300
hertz for signals in the model on the PSP,
the typical error per component is less
than 0.1%. The equivalent digital speed
required for this computation would need to
be about 970 million NOPS.
A fully expanded SIMSTAR multi- |
processor can include two PSFs operating in
parallel plus a Vector Frocesscr and the
Digital Arithmetic Processor. Also, if
needed the Host digital processors can
contribute to the simulation task. There-
fore, in total, the system performance at
this speed would be approximately 200
million NOFS.
DIGITAL ARITHMETIC PROCESSOR
Sequential processing in SIMSTAR is
performed by the Digital Arithmetic
Processor (DAP) which provides up to Il
million NOPS performance. As shown in the
block diagram (Figure 7), the DAF is built
around a System Bus which has a bandwidth
of over 6 million 32 bit words per second.
A sophisticated 32 bit CPU is available on
this bus. This CPU has firmware for
floating-point arithmetic in both single
and double precision. A basic Integrated
Memory Module (IMM) containing 1 megabyte
Qf 600 nanosecond MOS memory must be
connected to the bus. This is further
expandable within a SIMSTAR to a maximum of
2 megabytes. The addition of the optional
Floating Point Accelerator (FPA) increases
the speed of the DAP from 430 K Whetstones
to 668 KF Whetstones for the single
precision Whetstone benchmark. For the
double precision Whetstone benchmark, the
equivalent speeds are 204 K Whetstones
ree — oe eae aoe
| INTEGRATED | INTEGRATED 42 BIT a FLOATING |
MEMORY MEMORY 2 ) POINT |
p MODULE I MODULE | B ACCELE eligi
vo INTERRUPT F cogDATA 7 aeony
MER NV }
PROCESSOR CONTROL | ‘PROCESSOR CONTROLS
ae Tom |
et a 3°9-¢
) a
CONTROL
, PSP
FLOPPY ;
DISC LCP
Figure 7 — Digital Arithmetic Processor
(*
without the FPA and 465 K Whetstones with
the FPA. A single precision floating point
add is performed in 1.65 microseconds,
while a single precision flaating point
multiply is performed in 2.25 microseconds.
I/O devices are shown attached to the
bottom of the System Bus. The basic DAP is
operated through an Input/Output Processor,
which translates the System Bus to a simple
16 bit 1/0 Bus. This Bus controls the
floppy disk which is used to boot up the
SIMSTARK Operating System in the DAF. A
user can add a console CRY to the [7/0
Processor to provide local operation of the
SIMSTAR multiprocessor. The rest of the
devices attached to the System Bus are
interfaces to the Parallel Simulation
Frocessor in SIMSTAR. The Interrupt/Timer
Control provides a programmable, pricrity
interrupt capability into the CPU as well
as a variety of timers for controlling the
DAF and Data Conversion Processor
Operation. The Data Conversion Frocessor
(DCP) is an optional high speed data
conversion system to communicate with one
Or two parallel simulation processors.
This 18 an intelligent, programmable device
im which up to 6 tasks may be activated
Simultaneously to convert continuous
Signals to floating-point data. Also, wix
additional tasks can be activated to
convert floating-point data to continuous
Signals. In a SIMSTAR system with two
PSPs, half the DCP tasks are controlled by
each PLU. |
The basic interface between the DAP
and the PSP is through the Remote Memory
Controls (RMC) which are a memory mapped
means of accessing the various data devices
in the PSF and the LCP. The RMC device is
setup to consume part of the extended
memory addressing range of the CPU so that
programs may simply store floating-point
data into specified memory locations and
the LCP will perform the necessary oper—
ations to control or setup the PSP. As was
previously described in the PSP section,
the Host computer will also communicate
through the SIMBus in the PSP so that it
can down-load into the DAF local memory.
AN appropriate executive is provided with
SIMSTAR to take a block of data/program at
&@ time through this Remote Memory Control
for transfer to the appropriate local
memory of the DAF.
PUNCTION GENERATION PROCESSOR
One of the most important functional
requirements in simulation applications of
the SIMSTAR multiprocessor is the repre-
sentation of empirical data such as that
derived from wind tunnel tests of serospace
vehicles, steam tables for nuclear reactor
Simulations, and compressor maps for
turbine engine simulation. Many types of
special-purpose devices and computer pro-
grams have been developed to implement this
requirement. In SIMSTAR, single variable
function generation devices are incor-
porated in the PMU and an efficient soft~
ware package is provided for the DAP to
represent multi-variable functions. I[n
addition, one can add a Function Generation
Processor (FGP) which extends the perform
ance of SIMSTAR for this requirement by up
to 2 decades of speed.
f(X, Y)
Figure 8 — A Function of Two Variables
A function of 2 variables, as shown
graphically in Figure 8, is a surface which
can be defined by a set of ordered triples
( F, X, Y ). Usually, the data is aligned
on a rectangular grid so that the function
values may be stored in a two dimensional
array F( I, dg ) in which the subscripts
represent grid values of the independent
variables. Function generation is a table
look-up and interpolation process; that is,
given an instantaneous value of X and Y,
the computing device must compute an
address in the two dimensional array to
pick up the appropriate four adjacent
points and perform the necessary inter
polation on that “surface” to approximate
the original continuous function. This
requirement extends to functions of 3
variables to interpolate between surfaces
and up to 6 variables in some applications.
In a typical data set for an aerospace
vehicle, a mixture of 50-100 multi-variable
Basic SIMSTAR
DSFG Fi DIGITALLY SET FUNCTION GENERATORS
Continuous function of one Variable with
45 or 221 equally spaced breakpoints.
A total of up to 24 DSFGs are provided
in each Parallel Mathematical Unit.
— re A ele A ee eh ER alll
pee es see ei eS pd ee ee ee ee es ge ee cee ee ee ces de SS ee SE ae eh A ree
FGESYS F1,2,3,4 FUNCTION GENERATION SYSTEM
Specialized software package for DAP to
compute discrete floating-point function
yalues at high speed. Compatible with
DCP for use with PSP. Also operates on
Host.
Figure 9a - Basic SIMSTAR Function
Generation
Digitally Set Function Generators (DSFG)
which provide continuous generation of
functions of 1 variable. The breakpoints
can be selected for either 45 or 271
equally spaced values between -1.1 and
+i.1. These units are setup through the
LCP and operate completely independent of
any other processor during the problem
solution. — .
For the representation of
multi-variable functions, the FGSYS -
Function Generation System software package —
is available for the Digital Arithmetic |
Processor (or the Host data processing
system). This package uses a unique code
generation process for a high level of user
convenience while obtaining optimum code
efficiency at run-time. The internal
operations of the run-time code are auto-
matically optimized to eliminate redundant
processing for common function argument
sets.
— 2 ee ee LLL ele ea ME SY eal reir ede
MULTIVARIABLE FUNCTION GENERATORS
Continuous function of twe to four
variables. Setup by DAP and run
independent.
VPFGSYS F1,2,3,4 &(5) VECTOR PROCESSOR FUNCTION
GENERATION SYSTEM
High speed floating-point arr ay
processor and software operating froe
DAP. Compatible with Data Conversian
Processor for use with PSP.
Figure 9b —- Elements of the Function
Generation Processor
An optional Function Generation
Processor can be added to SIMSTAR as
mentioned in the previous architecture
section. This sub-system is composed of
two different elements as shown in Figure
96. First, the specialized Multi-Variable
Function Generation (MVFG) subsystem is _
available which provides continuous func-
tion generation for 2, 3, or 4 variables.
Second, a Vector Processor Function
Generation System (VPFGSYS) is available
for floating-point function lookup and
interpolation at about 10 times the speed
of the DAP. This unit operates independent
Of the DAP once it is setup and can
communicate directly with the Data
Conversion Processor.
A block diagram of a function of 2
variables implemented in the MVFG is shown
an Figure 10. The continuous signals (X,Y)
Produced by Mathematical Computing Blocks
in the PMU are connected by the switch
matrix to the X and Y Input Normalizers.
X ADOAESS
WEIGHTING
TERME
V4
:
Y¥ ADORESS
Figure 10 -— FC(X,Y) Block in MVFG
These elements of the MVFG transform the
continuous signal into an Address and a
continuous increment as defined by the
breakpoint data storage K and Yi° The
continuous increments, A X and AY go toa
the Weighting Terms computing subsystem to
produce the continuous signals W 4 ~ W 4
.which feed the four Digital~-to-Analog
Multipliers (DAMs). Each of the DAMs has a
portion of the function storage allocated
to it and the appropriate cells are select-
ed by the X address and the Y address. The
four DAM outputs are summed in the final
output device to produce the continuous
signal *(X,Y). This signal is then fed to
the switch matrix of the PMU to be distrib-
uted to the appropriate Mathematical
Computing Blocks.
A compiete MVFG subsystem will
typically be composed of up to 6 pairs of
these continuous function generation units
to compute 12 £(X,Y)’s. In addition to
functions of 2 variables as described here,
the MVFG module can be re-configured auto-
matically so that a full MVFG unit can
provide 6 functions of three variables or 3$
functions of four variables. For a func-
tion of four variables, up to 80946 words of
function data storage are available. The
Signals which act as arguments to the MVFG
can contain computing frequencies up toa i!
KHz with accurate computation of the
function. }
For large sets of multi-variable
functions with high frequency arguments,
the VPFGSYS is available for SIMSTAR. This
is a Vector Processor based hardware /
software subsystem which is compatible with
the DAP on a Common Memory interface. This
subsystem is programmed in FORTRAN on the
Host using a standard library supplied by
EAI. The function data and control array
are down-loaded into the Vector Processor
by the SIMSTAR Setup program. When the
run-time task is activated in the DAP, it
commands the Vector Processor to produce
the specified functions from instantaneous
argument values produced either by the DAP
program or, through the DCP, from sampled.
continuous signals in the Paraliel Math
Unit. |
te ee ee re os
SIMSTAR programming system is the STARTRAN
programming language, which is an extended
version of the Continuous System Simulation
language standard (2). This language
allows the user to specify differential
equations in a natural form which 165 a
superset of FORTRAN. For SIMSTAR, the user
need only add a set of declarations to
identify the variables to be produced on
the FSP as distinguished from the DAF.
first stage of STARTRAN performs the
necessary syntax analysis and partitions
the problem into a sequential and a
parallel part. For processing the
sequential part. a high level language
Processor (D-TRAN) is provided to translate
the differential equations into a FORTRAN
program. This process also includes the
generation of the control regions of the
simulation for initialization and post-run
processing.
The
The parallel region of the model is
translated by the STARTRAN front-end into a
set of simulation language statements for
P-TRAN, which handles the reduction of
these equations into a parallel processing
structure for the mathematical computing
blocks of SIMSTAR. The output of P~TRAN is
a language for representing the connections
between these mathematical blocks and
logical operators which will be allocated
to the Farallel Math Unit and the Parallel
logical Unit of the PSP. Also, FP-TRAN
produces) a FORTRAN program including all of
the necessary data to setup the math-
ematical blocks of the PSP as well as the
FGF. |
The combination of the RUN, SETUF, and
FLOW files represent a low-level entry to
the SIMSTAR system. That 15, users can
prepare these routines by hand and then
process the files to setup and run the
SIMSTAR system. For instance, in FProblem—
Oriented Languages, it is often easier to
produce this lower level input than to try
to produce the higher level for STARTRAN.
Also, some efficiencies in the use of the
Various processors may be possible if the
user optimizes the program entered at this
level.
The digital RUN portion for the DAP is
compiled by FORTRAN, combined with various
libraries provided with SIMSTAR, and cat-—
aloged to create a Run-time task for the
DAP. Also, the SETUP file is compiled by
FORTRAN and cataloged using other libraries
to build a separate task. The FLOW file
that defines the topological connection
between the mathematical and logical com—
puting elements is processed through the
FP-~COMP compiler to produce the necessary
matrix images and binary patterns to
operate on the PMU and FLU.
The user runs this combined simulation
at oa terminal on the Host using the "Run-
Time Executive" which operates in the DAP.
This executive first activates the setup
task to load the Function Generation Fro~
cessor, the Parallel Mathematical Unit, and
the Farallel Logic Unit for the specific
data of the problem. This also includes
calls to routines to input the files cre~
ated by F-COMF. Then, the user can inin~
A es 6 et ee,
tiate runs using a standard test case from
the Run-time Exe