/[MITgcm]/manual/s_software/text/sarch.tex

Diff of /manual/s_software/text/sarch.tex

Parent Directory | Revision Log | View Revision Graph Revision Graph | View Patch Patch

-revision 1.5 by cnh,
Tue Nov 13 18:32:33 2001 UTC
+revision 1.12 by afe,
Wed Jan 28 19:59:03 2004 UTC
 Line 1
  % $Header$
- In this chapter we describe the software architecture and
+ This chapter focuses on describing the {\bf WRAPPER} environment within which
- implementation strategy for the MITgcm code. The first part of this
+ both the core numerics and the pluggable packages operate. The description
- chapter discusses the MITgcm architecture at an abstract level. In the second
+ presented here is intended to be a detailed exposition and contains significant
- part of the chapter we described practical details of the MITgcm implementation
+ background material, as well as advanced details on working with the WRAPPER.
- and of current tools and operating system features that are employed.
+ The tutorial sections of this manual (see sections
+ \ref{sect:tutorials}  and \ref{sect:tutorialIII})
+ contain more succinct, step-by-step instructions on running basic numerical
+ experiments, of varous types, both sequentially and in parallel. For many
+ projects simply starting from an example code and adapting it to suit a
+ particular situation
+ will be all that is required.
+ The first part of this chapter discusses the MITgcm architecture at an
+ abstract level. In the second part of the chapter we described practical
+ details of the MITgcm implementation and of current tools and operating system
+ features that are employed.
  \section{Overall architectural goals}
-Line 28 
 of
+Line 38 
 of
  \begin{enumerate}
  \item A core set of numerical and support code. This is discussed in detail in
- section \ref{sec:partII}.
+ section \ref{sect:partII}.
  \item A scheme for supporting optional "pluggable" {\bf packages} (containing
  for example mixed-layer schemes, biogeochemical schemes, atmospheric physics).
  These packages are used both to overlay alternate dynamics and to introduce
-Line 74 
 Environment Resource). All numerical and
+Line 84 
 Environment Resource). All numerical and
  to ``fit'' within the WRAPPER infrastructure. Writing code to ``fit'' within
  the WRAPPER means that coding has to follow certain, relatively
  straightforward, rules and conventions ( these are discussed further in
- section \ref{sec:specifying_a_decomposition} ).
+ section \ref{sect:specifying_a_decomposition} ).
  The approach taken by the WRAPPER is illustrated in figure
  \ref{fig:fit_in_wrapper} which shows how the WRAPPER serves to insulate code
-Line 87 
 and operating systems. This allows numer
+Line 97 
 and operating systems. This allows numer
  \resizebox{!}{4.5in}{\includegraphics{part4/fit_in_wrapper.eps}}
  \end{center}
  \caption{
- Numerical code is written too fit within a software support
+ Numerical code is written to fit within a software support
  infrastructure called WRAPPER. The WRAPPER is portable and
  can be specialized for a wide range of specific target hardware and
  programming environments, without impacting numerical code that fits
-Line 98 
 optimized for that platform.}
+Line 108 
 optimized for that platform.}
  \end{figure}
  \subsection{Target hardware}
- \label{sec:target_hardware}
+ \label{sect:target_hardware}
  The WRAPPER is designed to target as broad as possible a range of computer
  systems. The original development of the WRAPPER took place on a
-Line 110 
 uniprocessor and multi-processor Sun sys
+Line 120 
 uniprocessor and multi-processor Sun sys
  (UMA) and non-uniform memory access (NUMA) designs. Significant work has also
  been undertaken on x86 cluster systems, Alpha processor based clustered SMP
  systems, and on cache-coherent NUMA (CC-NUMA) systems from Silicon Graphics.
- The MITgcm code, operating within the WRAPPER, is also used routinely used on
+ The MITgcm code, operating within the WRAPPER, is also routinely used on
  large scale MPP systems (for example T3E systems and IBM SP systems). In all
  cases numerical code, operating within the WRAPPER, performs and scales very
  competitively with equivalent numerical code that has been modified to contain
-Line 118 
 native optimizations for a particular sy
+Line 128 
 native optimizations for a particular sy
  \subsection{Supporting hardware neutrality}
- The different systems listed in section \ref{sec:target_hardware} can be
+ The different systems listed in section \ref{sect:target_hardware} can be
  categorized in many different ways. For example, one common distinction is
  between shared-memory parallel systems (SMP's, PVP's) and distributed memory
  parallel systems (for example x86 clusters and large MPP systems). This is one
-Line 211 
 computational phases a processor will re
+Line 221 
 computational phases a processor will re
  whenever it requires values that outside the domain it owns. Periodically
  processors will make calls to WRAPPER functions to communicate data between
  tiles, in order to keep the overlap regions up to date (see section
- \ref{sec:communication_primitives}). The WRAPPER functions can use a
+ \ref{sect:communication_primitives}). The WRAPPER functions can use a
  variety of different mechanisms to communicate data between tiles.
  \begin{figure}
-Line 298 
 value to be communicated between CPU's.
+Line 308 
 value to be communicated between CPU's.
  \end{figure}
  \subsection{Shared memory communication}
- \label{sec:shared_memory_communication}
+ \label{sect:shared_memory_communication}
  Under shared communication independent CPU's are operating
  on the exact same global address space at the application level.
-Line 324 
 the systems main-memory interconnect. Th
+Line 334 
 the systems main-memory interconnect. Th
  communication very efficient provided it is used appropriately.
  \subsubsection{Memory consistency}
- \label{sec:memory_consistency}
+ \label{sect:memory_consistency}
  When using shared memory communication between
  multiple processors the WRAPPER level shields user applications from
-Line 348 
 memory, the WRAPPER provides a place to
+Line 358 
 memory, the WRAPPER provides a place to
  ensure memory consistency for a particular platform.
  \subsubsection{Cache effects and false sharing}
- \label{sec:cache_effects_and_false_sharing}
+ \label{sect:cache_effects_and_false_sharing}
  Shared-memory machines often have local to processor memory caches
  which contain mirrored copies of main memory. Automatic cache-coherence
-Line 367 
 in an application are potentially visibl
+Line 377 
 in an application are potentially visibl
  threads operating within a single process is the standard mechanism for
  supporting shared memory that the WRAPPER utilizes. Configuring and launching
  code to run in multi-threaded mode on specific platforms is discussed in
- section \ref{sec:running_with_threads}.  However, on many systems, potentially
+ section \ref{sect:running_with_threads}.  However, on many systems, potentially
  very efficient mechanisms for using shared memory communication between
  multiple processes (in contrast to multiple threads within a single
  process) also exist. In most cases this works by making a limited region of
-Line 380 
 distributed with the default WRAPPER sou
+Line 390 
 distributed with the default WRAPPER sou
  nature.
  \subsection{Distributed memory communication}
- \label{sec:distributed_memory_communication}
+ \label{sect:distributed_memory_communication}
  Many parallel systems are not constructed in a way where it is
  possible or practical for an application to use shared memory
  for communication. For example cluster systems consist of individual computers
-Line 394 
 described in \ref{hoe-hill:99} substitut
+Line 404 
 described in \ref{hoe-hill:99} substitut
  highly optimized library.
  \subsection{Communication primitives}
- \label{sec:communication_primitives}
+ \label{sect:communication_primitives}
  \begin{figure}
  \begin{center}
-Line 538 
 WRAPPER are
+Line 548 
 WRAPPER are
  computing CPU's.
  \end{enumerate}
  This section describes the details of each of these operations.
- Section \ref{sec:specifying_a_decomposition} explains how the way in which
+ Section \ref{sect:specifying_a_decomposition} explains how the way in which
  a domain is decomposed (or composed) is expressed. Section
- \ref{sec:starting_a_code} describes practical details of running codes
+ \ref{sect:starting_a_code} describes practical details of running codes
  in various different parallel modes on contemporary computer systems.
- Section \ref{sec:controlling_communication} explains the internal information
+ Section \ref{sect:controlling_communication} explains the internal information
  that the WRAPPER uses to control how information is communicated between
  tiles.
  \subsection{Specifying a domain decomposition}
- \label{sec:specifying_a_decomposition}
+ \label{sect:specifying_a_decomposition}
  At its heart much of the WRAPPER works only in terms of a collection of tiles
  which are interconnected to each other. This is also true of application
-Line 599 
 be created within a single process. Each
+Line 609 
 be created within a single process. Each
  dimensions of {\em sNx} and {\em sNy}. If, when the code is executed, these tiles are
  allocated to different threads of a process that are then bound to
  different physical processors ( see the multi-threaded
- execution discussion in section \ref{sec:starting_the_code} ) then
+ execution discussion in section \ref{sect:starting_the_code} ) then
  computation will be performed concurrently on each tile. However, it is also
  possible to run the same decomposition within a process running a single thread on
  a single processor. In this case the tiles will be computed over sequentially.
-Line 651 
 Within a {\em bi}, {\em bj} loop
+Line 661 
 Within a {\em bi}, {\em bj} loop
  computation is performed concurrently over as many processes and threads
  as there are physical processors available to compute.
+ An exception to the the use of {\em bi} and {\em bj} in loops arises in the
+ exchange routines used when the exch2 package is used with the cubed
+ sphere.  In this case {\em bj} is generally set to 1 and the loop runs from
+,{\em bi}.  Within the loop {\em bi} is used to retrieve the tile number,
+ which is then used to reference exchange parameters.
  The amount of computation that can be embedded
  a single loop over {\em bi} and {\em bj} varies for different parts of the
  MITgcm algorithm. Figure \ref{fig:bibj_extract} shows a code extract
-Line 771 
 The global domain size is again ninety g
+Line 787 
 The global domain size is again ninety g
  forty grid points in y. The two sub-domains in each process will be computed
  sequentially if they are given to a single thread within a single process.
  Alternatively if the code is invoked with multiple threads per process
- the two domains in y may be computed on concurrently.
+ the two domains in y may be computed concurrently.
  \item
  \begin{verbatim}
        PARAMETER (
-Line 790 
 There are six tiles allocated to six sep
+Line 806 
 There are six tiles allocated to six sep
  This set of values can be used for a cube sphere calculation.
  Each tile of size $32 \times 32$ represents a face of the
  cube. Initializing the tile connectivity correctly ( see section
- \ref{sec:cube_sphere_communication}. allows the rotations associated with
+ \ref{sect:cube_sphere_communication}. allows the rotations associated with
  moving between the six cube faces to be embedded within the
  tile-tile communication code.
  \end{enumerate}
  \subsection{Starting the code}
- \label{sec:starting_the_code}
+ \label{sect:starting_the_code}
  When code is started under the WRAPPER, execution begins in a main routine {\em
  eesupp/src/main.F} that is owned by the WRAPPER. Control is transferred
  to the application through a routine called {\em THE\_MODEL\_MAIN()}
-Line 807 
 by the application code. The startup cal
+Line 823 
 by the application code. The startup cal
  WRAPPER is shown in figure \ref{fig:wrapper_startup}.
  \begin{figure}
+ {\footnotesize
  \begin{verbatim}
         MAIN
-Line 835 
 WRAPPER is shown in figure \ref{fig:wrap
+Line 852 
 WRAPPER is shown in figure \ref{fig:wrap
  \end{verbatim}
+ }
  \caption{Main stages of the WRAPPER startup procedure.
  This process proceeds transfer of control to application code, which
  occurs through the procedure {\em THE\_MODEL\_MAIN()}.
-Line 842 
 occurs through the procedure {\em THE\_M
+Line 860 
 occurs through the procedure {\em THE\_M
  \end{figure}
  \subsubsection{Multi-threaded execution}
- \label{sec:multi-threaded-execution}
+ \label{sect:multi-threaded-execution}
  Prior to transferring control to the procedure {\em THE\_MODEL\_MAIN()} the
  WRAPPER may cause several coarse grain threads to be initialized. The routine
  {\em THE\_MODEL\_MAIN()} is called once for each thread and is passed a single
  stack argument which is the thread number, stored in the
  variable {\em myThid}. In addition to specifying a decomposition with
- multiple tiles per process ( see section \ref{sec:specifying_a_decomposition})
+ multiple tiles per process ( see section \ref{sect:specifying_a_decomposition})
  configuring and starting a code to run using multiple threads requires the following
  steps.\\
-Line 917 
 File: {\em eesupp/inc/MAIN\_PDIRECTIVES1
+Line 935 
 File: {\em eesupp/inc/MAIN\_PDIRECTIVES1
  File: {\em eesupp/inc/MAIN\_PDIRECTIVES2.h}\\
  File: {\em model/src/THE\_MODEL\_MAIN.F}\\
  File: {\em eesupp/src/MAIN.F}\\
- File: {\em tools/genmake}\\
+ File: {\em tools/genmake2}\\
  File: {\em eedata}\\
  CPP:  {\em TARGET\_SUN}\\
  CPP:  {\em TARGET\_DEC}\\
-Line 930 
 Parameter:  {\em nTy}
+Line 948 
 Parameter:  {\em nTy}
  } \\
  \subsubsection{Multi-process execution}
- \label{sec:multi-process-execution}
+ \label{sect:multi-process-execution}
  Despite its appealing programming model, multi-threaded execution remains
  less common then multi-process execution. One major reason for this
-Line 942 
 models varies between systems.
+Line 960 
 models varies between systems.
  Multi-process execution is more ubiquitous.
  In order to run code in a multi-process configuration a decomposition
- specification ( see section \ref{sec:specifying_a_decomposition})
+ specification ( see section \ref{sect:specifying_a_decomposition})
  is given ( in which the at least one of the
  parameters {\em nPx} or {\em nPy} will be greater than one)
  and then, as for multi-threaded operation,
-Line 956 
 critical communication. However, in orde
+Line 974 
 critical communication. However, in orde
  of controlling and coordinating the start up of a large number
  (hundreds and possibly even thousands) of copies of the same
  program, MPI is used. The calls to the MPI multi-process startup
- routines must be activated at compile time. This is done
+ routines must be activated at compile time.  Currently MPI libraries are
- by setting the {\em ALLOW\_USE\_MPI} and {\em ALWAYS\_USE\_MPI}
+ invoked by
- flags in the {\em CPP\_EEOPTIONS.h} file.\\
+ specifying the appropriate options file with the
+ \begin{verbatim}-of\end{verbatim} flag when running the {\em genmake2}
- \fbox{
+ script, which generates the Makefile for compiling and linking MITgcm.
- \begin{minipage}{4.75in}
+ (Previously this was done by setting the {\em ALLOW\_USE\_MPI} and
- File: {\em eesupp/inc/CPP\_EEOPTIONS.h}\\
+ {\em ALWAYS\_USE\_MPI} flags in the {\em CPP\_EEOPTIONS.h} file.)  More
- CPP:  {\em ALLOW\_USE\_MPI}\\
+ detailed information about the use of {\em genmake2} for specifying
- CPP:  {\em ALWAYS\_USE\_MPI}\\
+ local compiler flags is located in section 3 ??\\
- Parameter:  {\em nPx}\\
- Parameter:  {\em nPy}
- \end{minipage}
- } \\
- Additionally, compile time options are required to link in the
- MPI libraries and header files. Examples of these options
- can be found in the {\em genmake} script that creates makefiles
- for compilation. When this script is executed with the {bf -mpi}
- flag it will generate a makefile that includes
- paths for search for MPI head files and for linking in
- MPI libraries. For example the {\bf -mpi} flag on a
-  Silicon Graphics IRIX system causes a
- Makefile with the compilation command
- Graphics IRIX system \begin{verbatim}
- mpif77 -I/usr/local/mpi/include -DALLOW_USE_MPI -DALWAYS_USE_MPI
- \end{verbatim}
- to be generated.
- This is the correct set of options for using the MPICH open-source
- version of MPI, when it has been installed under the subdirectory
- /usr/local/mpi.
- However, on many systems there may be several
- versions of MPI installed. For example many systems have both
- the open source MPICH set of libraries and a vendor specific native form
- of the MPI libraries. The correct setup to use will depend on the
- local configuration of your system.\\
  \fbox{
  \begin{minipage}{4.75in}
- File: {\em tools/genmake}
+ File: {\em tools/genmake2}
  \end{minipage}
  } \\
  \paragraph{\bf Execution} The mechanics of starting a program in
-Line 1112 
 A value of {\em COMM\_NONE} is used to i
+Line 1105 
 A value of {\em COMM\_NONE} is used to i
  neighbor to communicate with on a particular face. A value
  of {\em COMM\_MSG} is used to indicated that some form of distributed
  memory communication is required to communicate between
- these tile faces ( see section \ref{sec:distributed_memory_communication}).
+ these tile faces ( see section \ref{sect:distributed_memory_communication}).
  A value of {\em COMM\_PUT} or {\em COMM\_GET} is used to indicate
  forms of shared memory communication ( see section
- \ref{sec:shared_memory_communication}). The {\em COMM\_PUT} value indicates
+ \ref{sect:shared_memory_communication}). The {\em COMM\_PUT} value indicates
  that a CPU should communicate by writing to data structures owned by another
  CPU. A {\em COMM\_GET} value indicates that a CPU should communicate by reading
  from data structures owned by another CPU. These flags affect the behavior
-Line 1166 
 the product of the parameters {\em nTx}
+Line 1159 
 the product of the parameters {\em nTx}
  are read from the file {\em eedata}. If the value of {\em nThreads}
  is inconsistent with the number of threads requested from the
  operating system (for example by using an environment
- variable as described in section \ref{sec:multi_threaded_execution})
+ variable as described in section \ref{sect:multi_threaded_execution})
  then usually an error will be reported by the routine
  {\em CHECK\_THREADS}.\\
-Line 1184 
 Parameter: {\em nTy} \\
+Line 1177 
 Parameter: {\em nTy} \\
  }
  \item {\bf memsync flags}
- As discussed in section \ref{sec:memory_consistency}, when using shared memory,
+ As discussed in section \ref{sect:memory_consistency}, when using shared memory,
  a low-level system function may be need to force memory consistency.
  The routine {\em MEMSYNC()} is used for this purpose. This routine should
  not need modifying and the information below is only provided for
-Line 1210 
 asm("lock; addl $0,0(%%esp)": : :"memory
+Line 1203 
 asm("lock; addl $0,0(%%esp)": : :"memory
  \end{verbatim}
  \item {\bf Cache line size}
- As discussed in section \ref{sec:cache_effects_and_false_sharing},
+ As discussed in section \ref{sect:cache_effects_and_false_sharing},
  milti-threaded codes explicitly avoid penalties associated with excessive
  coherence traffic on an SMP system. To do this the shared memory data structures
  used by the {\em GLOBAL\_SUM}, {\em GLOBAL\_MAX} and {\em BARRIER} routines
-Line 1238 
 or {\em GLOBAL\_SUM\_R4()} (for 32-bit f
+Line 1231 
 or {\em GLOBAL\_SUM\_R4()} (for 32-bit f
  setting for the \_GSUM macro is given in the file {\em CPP\_EEMACROS.h}.
  The \_GSUM macro is a performance critical operation, especially for
  large processor count, small tile size configurations.
- The custom communication example discussed in section \ref{sec:jam_example}
+ The custom communication example discussed in section \ref{sect:jam_example}
  shows how the macro is used to invoke a custom global sum routine
  for a specific set of hardware.
-Line 1252 
 physical fields and whether fields are 3
+Line 1245 
 physical fields and whether fields are 3
  in the header file {\em CPP\_EEMACROS.h}. As with \_GSUM, the
  \_EXCH operation plays a crucial role in scaling to small tile,
  large logical and physical processor count configurations.
- The example in section \ref{sec:jam_example} discusses defining an
+ The example in section \ref{sect:jam_example} discusses defining an
  optimized and specialized form on the \_EXCH operation.
  The \_EXCH operation is also central to supporting grids such as
-Line 1292 
 This can be achieved using a Fortran 90
+Line 1285 
 This can be achieved using a Fortran 90
  if this might be unavailable then the work arrays can be extended
  with dimensions use the tile dimensioning scheme of {\em nSx}
  and {\em nSy} ( as described in section
- \ref{sec:specifying_a_decomposition}). However, if the configuration
+ \ref{sect:specifying_a_decomposition}). However, if the configuration
  being specified involves many more tiles than OS threads then
  it can save memory resources to reduce the variable
  {\em MAX\_NO\_THREADS} to be equal to the actual number of threads that
-Line 1351 
 Here we show how it can be used to impro
+Line 1344 
 Here we show how it can be used to impro
  how it can be used to adapt to new griding approaches.
  \subsubsection{JAM example}
- \label{sec:jam_example}
+ \label{sect:jam_example}
  On some platforms a big performance boost can be obtained by
  binding the communication routines {\em \_EXCH} and
  {\em \_GSUM} to specialized native libraries ) fro example the
-Line 1374 
 Developing specialized code for other li
+Line 1367 
 Developing specialized code for other li
  pattern.
  \subsubsection{Cube sphere communication}
- \label{sec:cube_sphere_communication}
+ \label{sect:cube_sphere_communication}
  Actual {\em \_EXCH} routine code is generated automatically from
  a series of template files, for example {\em exch\_rx.template}.
  This is done to allow a large number of variations on the exchange
-Line 1407 
 quantities at the C-grid vorticity point
+Line 1400 
 quantities at the C-grid vorticity point
  Fitting together the WRAPPER elements, package elements and
  MITgcm core equation elements of the source code produces calling
- sequence shown in section \ref{sec:calling_sequence}
+ sequence shown in section \ref{sect:calling_sequence}
  \subsection{Annotated call tree for MITgcm and WRAPPER}
- \label{sec:calling_sequence}
+ \label{sect:calling_sequence}
  WRAPPER layer.
+ {\footnotesize
  \begin{verbatim}
         MAIN
-Line 1441 
 WRAPPER layer.
+Line 1435 
 WRAPPER layer.
         |--THE_MODEL_MAIN   :: Numerical code top-level driver routine
  \end{verbatim}
+ }
  Core equations plus packages.
+ {\footnotesize
  \begin{verbatim}
  C
  C
-Line 1782 
 C    |-COMM_STATS     :: Summarise inter
+Line 1778 
 C    |-COMM_STATS     :: Summarise inter
  C                     :: events.
  C
  \end{verbatim}
+ }
  \subsection{Measuring and Characterizing Performance}

 Legend:



Removed from v.1.5
 


changed lines


 
Added in v.1.12
 Legend:



Removed from v.1.5
 


changed lines


 
Added in v.1.12
-Removed from v.1.5
+Added in v.1.12

	ViewVC Help
Powered by ViewVC 1.1.22