Opens a larger view. Escape closes it.

hardware-counters

hpcg_c16_profile.txt

CrayPat/X:  Version 23.09.0 Revision 6034a9414 sles15.4_x86_64  08/01/23 18:55:19

Number of PEs (MPI ranks):   16
                           
Numbers of PEs per Node:     16
                           
Numbers of Threads per PE:    1
                           
Number of Cores per Socket:  64

Execution start time:  Tue Jul 28 08:40:59 2026

System name and speed:  nid002767  2.212 GHz (nominal)

AMD   Rome                 CPU  Family: 23  Model: 49  Stepping:  0

Core Performance Boost:  All 16 PEs have CPB capability

Current path to data file:
  /work/project/project/user/runs/sweep/hpcg_c16//exp_A   (RTS)


Notes for table 1:

  This table shows functions that have significant exclusive time,
    averaged across ranks.
  For further explanation, see the "General table notes" below, or 
    use:  pat_report -v -O profile ...

Table 1:  Profile by Function Group and Function

  Time% |      Time |     Imb. |  Imb. |     Calls | Group
        |           |     Time | Time% |           |  Function
        |           |          |       |           |   PE=HIDE
       
 100.0% | 26.775097 |       -- |    -- | 224,555.0 | Total
|-----------------------------------------------------------------------------
|  85.8% | 22.981523 |       -- |    -- |  20,179.0 | USER
||----------------------------------------------------------------------------
||  57.4% | 15.362584 | 1.574053 |  9.9% |       1.0 | main
||  26.1% |  6.976604 | 0.315374 |  4.6% |   2,303.0 | ComputeSPMV_ref.LOOP@li.59
||============================================================================
|  10.2% |  2.742858 |       -- |    -- | 195,140.0 | MPI
||----------------------------------------------------------------------------
||   5.9% |  1.589831 | 0.641222 | 30.7% |  56,385.0 | MPI_Wait
||   3.7% |  0.988397 | 1.155492 | 57.5% |  56,385.0 | MPI_Send
||============================================================================
|   3.5% |  0.930271 |       -- |    -- |   1,766.0 | MPI_SYNC
||----------------------------------------------------------------------------
||   3.4% |  0.910737 | 0.811414 | 89.1% |   1,763.0 | MPI_Allreduce(sync)
|=============================================================================

========================  Additional details  ========================

General table notes:

    The default notes for a table are based on the default definition of
    the table, and do not account for the effects of command-line options
    that may modify the content of the table.
    
    Detailed notes, produced by the pat_report -v option, do account for
    all command-line options, and also show how data is aggregated, and
    if the table content is limited by thresholds, rank selections, etc.
    
    An imbalance metric in a line is based on values in main threads
    across multiple ranks, or on values across all threads, as applicable.
    
    An imbalance percent in a line is relative to the maximum value
    for that line across ranks or threads, as applicable.
    
    If the number of Calls for a function is shown as "--", then that
    function was not traced and the other values in its line summarize
    the data collected for functions that it calls and that were traced.
    
Experiment:  trace

Original path to data file:
  /mnt/lustre/a2fs-work3/work/project/project/user/runs/sweep/hpcg_c16/exp_A/xf-files   (RTS)

Original program:
  /mnt/lustre/a2fs-work3/work/project/project/user/builds/hpcg/xhpcg

Instrumented with:  pat_build -g mpi,io -w -o xhpcg+pat xhpcg

Instrumented program:
  /mnt/lustre/a2fs-work3/work/project/project/user/runs/sweep/hpcg_c16/./xhpcg+pat

Program invocation:
  /mnt/lustre/a2fs-work3/work/project/project/user/runs/sweep/hpcg_c16/./xhpcg+pat

Exit Status:  0 for 16 PEs

Memory pagesize:  4 KiB

Memory hugepagesize:  Not Available

Programming environment:  CRAY

Runtime environment variables:
  CRAYPAT_COMPILER_OPTIONS=1
  CRAYPAT_GCC_LIB_PATH=/opt/cray/pe/gcc-libs
  CRAYPAT_LD_LIBRARY_PATH=/opt/cray/pe/gcc-libs/opt/cray/pe/perftools/23.09.0/lib64
  CRAYPAT_OPTS_EXECUTABLE=libexec64/opts
  CRAYPAT_ROOT=/opt/cray/pe/perftools/23.09.0
  CRAYPE_VERSION=2.7.23
  CRAY_BINUTILS_VERSION=/opt/cray/pe/cce/16.0.1
  CRAY_CC_VERSION=16.0.1
  CRAY_DSMML_VERSION=0.2.2
  CRAY_FFTW_VERSION=3.3.10.5
  CRAY_FTN_VERSION=16.0.1
  CRAY_LIBSCI_VERSION=23.09.1.1
  CRAY_MPICH_VERSION=8.1.27
  CRAY_PERFTOOLS_VERSION=23.09.0
  FFTW_VERSION=3.3.10.5
  LIBSCI_VERSION=23.09.1.1
  LMOD_FAMILY_COMPILER_VERSION=16.0.1
  LMOD_FAMILY_CRAYPE_CPU_VERSION=false
  LMOD_FAMILY_CRAYPE_NETWORK_VERSION=false
  LMOD_FAMILY_CRAYPE_VERSION=2.7.23
  LMOD_FAMILY_LIBFABRIC_VERSION=1.12.1.2.2.0.0
  LMOD_FAMILY_MPI_VERSION=8.1.27
  LMOD_FAMILY_PERFTOOLS_VERSION=false
  LMOD_FAMILY_PRGENV_VERSION=8.4.0
  LMOD_VERSION=8.7.19
  MPICH_DIR=/opt/cray/pe/mpich/8.1.27/ofi/crayclang/14.0
  OMP_NUM_THREADS=1
  PAT_RT_EXPDIR_BASE=/work/project/project/user/runs/sweep/hpcg_c16
  PAT_RT_EXPDIR_NAME=exp_A
  PAT_RT_EXPDIR_REPLACE=1
  PAT_RT_PERFCTR=PAPI_FP_OPS,PAPI_FP_INS,PAPI_TOT_INS,PAPI_TOT_CYC,cray_rapl:::PACKAGE_ENERGY,cray_rapl:::PP0_ENERGY,cray_zenl3:::UNC_L3_CACHE_MISSES,cray_zenl3:::UNC_L3_CACHE_REQUESTS,cray_zenl3:::UNC_L3_MISS_LATENCY
  PAT_RT_PERFCTR_DISABLE_COMPONENTS=nvml,rocm_smi
  PERFTOOLS_VERSION=23.09.0
  PMI_CONTROL_PORT=29693
  PMI_JOBID=14534144
  PMI_LOCAL_RANK=0
  PMI_LOCAL_SIZE=16
  PMI_RANK=0
  PMI_SHARED_SECRET=CHANGE_ME
  PMI_SIZE=16
  PMI_UNIVERSE_SIZE=16

Report time environment variables:
    CRAYPAT_ROOT=/opt/cray/pe/perftools/23.09.0

Number of MPI control variables collected:  136

  (To see the list, specify: -s mpi_cvar=show)

Report command line options:  -O profile

Operating system:
  Linux 5.14.21-150400.24.228-default #1 SMP PREEMPT_DYNAMIC Tue Jul 7 09:49:36 UTC 2026 (9ba75fc)

Hardware performance counter events:
   PAPI_TOT_INS           Instructions completed
   PAPI_FP_INS            Floating point instructions
   PAPI_TOT_CYC           Total cycles
   PAPI_FP_OPS            Floating point operations
   UNC_L3_CACHE_REQUESTS  Requests to L3 cache
   UNC_L3_MISS_LATENCY    Accumulated L3 miss latency (total cycles for all transactions divided by 16)
   UNC_L3_CACHE_MISSES    L3 cache misses by request type
   PACKAGE_ENERGY         Energy used by chip package
   PP0_ENERGY             Energy used by all cores in package

Estimated minimum instrumentation overhead per call of a traced function,
  which was subtracted from the data shown in this report
  (for raw data, use the option:  -s overhead=include):
    Time  2.448  microsecs

Number of traced functions that were called:  21

  (To see the list, specify:  -s traced_functions=show)


Warnings:
OpenMP regions included 1 region with no end address, and
1 region with an invalid address range, and they were ignored.

An exit() call or a STOP statement can cause missing end addresses.
Invalid address ranges indicate a problem with data collection.