CPU2017 Result Flag Description

Base Optimization Flags

C benchmarks

- -m64
- CC, LD
- Generate code for a 64-bit environment. The 64-bit environment sets int to 32 bits and long and pointer to 64 bits and generates code for AMD's x86-64 architecture. The compiler generates AMD64, INTEL64, x86-64 64-bit ABI. The default on a 32-bit host is 32-bit ABI. The default on a 64-bit host is 64-bit ABI if the target platform specified is 64-bit, otherwise the default is 32-bit.
- -Wl,-allow-multiple-definition
- LDCFLAGS
- Do not generate an error when linking multiple symbols of the same name.
- -Wl,-mllvm -Wl,-enable-licm-vrp
- LDCFLAGS
- Enables estimation of the virtual register pressure before performing loop invariant code motion. This estimation is used to decide the invariants that will be hoisted during loop invariant code motion.
- -flto
- COPTIMIZE, EXTRA_LDFLAGS
- Generate output files in LLVM formats suitable for link time optimization. When used with -S this generates LLVM intermediate language assembly files, otherwise this generates LLVM bitcode format object files (which may be passed to the linker depending on the stage selection options).
- -Wl,-mllvm -Wl,-region-vectorize
- EXTRA_LDFLAGS
- This flag enables vectorization of loops with complex control flow that can not be vectorized by loop and slp vectorizers.
- -Wl,-mllvm -Wl,-function-specialize
- EXTRA_LDFLAGS
- This option enables an optimization that generates and calls specialized function versions when they are called with constant arguments. This optimization helps in function inlining.
- -Wl,-mllvm -Wl,-align-all-nofallthru-blocks=6
- EXTRA_LDFLAGS
- Force the alignment of all blocks that have no fall-through predecessors (i.e. don't add nops that are executed). In log2 format (e.g 4 means align on 16B boundaries).
- -Wl,-mllvm -Wl,-reduce-array-computations=3
- EXTRA_LDFLAGS
- This option eliminates the array computations based on their usage. The computations on unused array elements and computations on zero valued array elements are eliminated with this optimization. -flto as whole program analysis is required to perform this optimization.
  
  Possible values:
  - 1: Eliminates the computations on unused array elements
  - 2: Eliminates the computations on zero valued array elements
  - 3: Eliminates the computations on unused and zero valued array elements
- -O3
- COPTIMIZE
- Like -O2, except that it enables optimizations that take longer to perform or that may generate larger code (in an attempt to make the program run faster).
  
  If multiple "O" options are used, with or without level numbers, the last such option is the one that is effective.
- Includes:
  - -O2
    - -O1
- -ffast-math
- COPTIMIZE
- Enables a range of optimizations that provide faster, though sometimes less precise, mathematical operations that may not conform to the IEEE-754 specifications. When this option is specified, the __STDC_IEC_559__ macro is ignored even if set by the system headers.
- -march=znver3
- COPTIMIZE
- Specify that Clang should generate code for a specific processor family member and later. For example, if you specify -march=znver1, the compiler is allowed to generate instructions that are valid on AMD Zen processors, but which may not exist on earlier products.
- -fveclib=AMDLIBM
- COPTIMIZE
- Use the given vector functions library.
- -fstruct-layout=5
- COPTIMIZE
- Analyzes the whole program to determine if the structures in the code can be peeled and if pointer or integer fields in the structure can be compressed. If feasible, this optimization transforms the code to enable these improvements. This transformation is likely to improve cache utilization and memory bandwidth. This, in turn, is expected to improve the scalability of programs executed on multiple cores.
  
  This is effective only under -flto as whole program analysis is required to perform this optimization. You can choose different levels of aggressiveness with which this optimization can be applied to your application with 1 being the least aggressive and 7 being the most aggressive level.
  
  Possible values:
  - fstruct-layout=1: enables structure peeling.
  - fstruct-layout=2: enables structure peeling and selectively compresses self-referential pointers in these structures to 32-bit pointers wherever safe.
  - fstruct-layout=3: enables structure peeling and selectively compresses self-referential pointers in these structures to 16-bit pointers wherever safe.
  - fstruct-layout=4: enables structure peeling, pointer compression as in level 2 and further enables compression of structure fields which are of integer type. This is performed under a strict safety check.
  - fstruct-layout=5: enables structure peeling, pointer compression as in level 3 and further enables compression of structure fields which are of integer type. This is performed under a strict safety check.
  - fstruct-layout=6: enables structure peeling, pointer compression as in level 2 and further enables compression of structure fields which are of type 64-bit 'signed int' or 'unsigned Int'. The user needs to ensure that the values assigned to 64-bit 'signed int' fields are in range -(2^31 - 1) to +(2^31 - 1) and 64-bit 'unsigned int' fields are in range 0 to +(2^31 - 1), otherwise, incorrect results may be obtained. This compression is performed without considering any safety analysis and so the user needs to ensure the safety based on the program compiled.
  - fstruct-layout=7: enables structure peeling, pointer compression as in level 3 and further enables compression of structure fields which are of type 64-bit 'signed int' or 'unsigned Int'. The user needs to ensure that the values assigned to 64-bit 'signed int' fields are in range -(2^31 - 1) to +(2^31 - 1) and 64-bit 'unsigned int' field are in range 0 to +(2^31 - 1) , otherwise, incorrect results may be obtained. This compression is performed without considering any safety analysis and so the user needs to ensure the safety based on the program compiled.
  Note:
  fstruct-layout=4 and fstruct-layout=5 are derived from fstruct-layout=2 and fstruct-layout=3 respectively with the added feature of safe compression of integer fields in structures. Going from fstruct-layout=4 to fstruct-layout=5 may result in higher performance if the pointer values are such that the pointers can be compressed to 16-bits.
  
  fstruct-layout=6 and fstruct-layout=7 are derived from fstruct-layout=2 and fstruct-layout=3 respectively with the added feature of compression of integer fields in structures. These are similar to fstruct-layout=4 and fstruct-layout=5, but here, the integer fields of the structures are always compressed from 64-bits to 32-bits without any safety guarantee.
- -mllvm -unroll-threshold=50
- COPTIMIZE
- Sets the limit at which loops will be unrolled. For example, if unroll threshold is set to 100 then only loops with 100 or fewer instructions will be unrolled.
- -mllvm -inline-threshold=1000
- COPTIMIZE
- Sets the compiler's inlining threshold level to the value passed as the argument. The inline threshold is used in the inliner heuristics to decide which functions should be inlined.
- -fremap-arrays
- COPTIMIZE
- This option enables an optimization that transforms the data layout of a single dimensional array to provide better cache locality by analysing the access patterns.
- -mllvm -function-specialize
- COPTIMIZE
- This option enables an optimization that generates and calls specialized function versions when they are called with constant arguments. This optimization helps in function inlining.
- -flv-function-specialization
- COPTIMIZE
- This option enables an optimization that generates and calls specialized function versions when the loops inside function are vectorizable and the arguments are not aliased with each other. This optimization helps in function inlining and vectorization.
- -mllvm -enable-gvn-hoist
- COPTIMIZE
- This option enables the GVN hoist pass, which is used to hoist computations from branches.
- -mllvm -global-vectorize-slp=true
- COPTIMIZE
- This option enables an optimization that does the slp vectorization across basic blocks. The SLP vectorizer vectorizes instructions within basic blocks. The global slp vectorizer analyzes instructions across basic blocks and vectorizes them.
- -mllvm -enable-licm-vrp
- COPTIMIZE
- Enables estimation of the virtual register pressure before performing loop invariant code motion. This estimation is used to decide the invariants that will be hoisted during loop invariant code motion.
- -mllvm -reduce-array-computations=3
- COPTIMIZE
- This option eliminates the array computations based on their usage. The computations on unused array elements and computations on zero valued array elements are eliminated with this optimization. -flto as whole program analysis is required to perform this optimization.
  
  Possible values:
  - 1: Eliminates the computations on unused array elements
  - 2: Eliminates the computations on zero valued array elements
  - 3: Eliminates the computations on unused and zero valued array elements
- -z muldefs
- LDOPTIMIZE
- Instructs the linker to use the first definition encountered for a symbol, and ignore all others.
- -lamdlibm
- EXTRA_LIBS
- Instructs the compiler to link with AMD-supported optimized math library.
- -ljemalloc
- EXTRA_LIBS
- Use the jemalloc library, which is a general purpose malloc(3) implementation that emphasizes fragmentation avoidance and scalable concurrency support.
- -lflang
- EXTRA_LIBS
- Instructs the compiler to link with flang Fortran runtime libraries.
- -lflangrti
- EXTRA_LIBS
- Instructs the compiler to link with flang Fortran runtime libraries.

C++ benchmarks

- -m64
- CXX, LD
- Generate code for a 64-bit environment. The 64-bit environment sets int to 32 bits and long and pointer to 64 bits and generates code for AMD's x86-64 architecture. The compiler generates AMD64, INTEL64, x86-64 64-bit ABI. The default on a 32-bit host is 32-bit ABI. The default on a 64-bit host is 64-bit ABI if the target platform specified is 64-bit, otherwise the default is 32-bit.
- -std=c++98
- CXX
- Selects the C++ language dialect.
- -Wl,-mllvm -Wl,-do-block-reorder=aggressive
- LDCXXFLAGS
- Block Reordering
  
  Possible values:
  - none: No block reorder
  - simple: Simple block reorder
  - aggressive: Aggressive block reorder
- -flto
- CXXOPTIMIZE, EXTRA_LDFLAGS
- Generate output files in LLVM formats suitable for link time optimization. When used with -S this generates LLVM intermediate language assembly files, otherwise this generates LLVM bitcode format object files (which may be passed to the linker depending on the stage selection options).
- -Wl,-mllvm -Wl,-region-vectorize
- EXTRA_LDFLAGS
- This flag enables vectorization of loops with complex control flow that can not be vectorized by loop and slp vectorizers.
- -Wl,-mllvm -Wl,-function-specialize
- EXTRA_LDFLAGS
- This option enables an optimization that generates and calls specialized function versions when they are called with constant arguments. This optimization helps in function inlining.
- -Wl,-mllvm -Wl,-align-all-nofallthru-blocks=6
- EXTRA_LDFLAGS
- Force the alignment of all blocks that have no fall-through predecessors (i.e. don't add nops that are executed). In log2 format (e.g 4 means align on 16B boundaries).
- -Wl,-mllvm -Wl,-reduce-array-computations=3
- EXTRA_LDFLAGS
- This option eliminates the array computations based on their usage. The computations on unused array elements and computations on zero valued array elements are eliminated with this optimization. -flto as whole program analysis is required to perform this optimization.
  
  Possible values:
  - 1: Eliminates the computations on unused array elements
  - 2: Eliminates the computations on zero valued array elements
  - 3: Eliminates the computations on unused and zero valued array elements
- -O3
- CXXOPTIMIZE
- Like -O2, except that it enables optimizations that take longer to perform or that may generate larger code (in an attempt to make the program run faster).
  
  If multiple "O" options are used, with or without level numbers, the last such option is the one that is effective.
- Includes:
  - -O2
    - -O1
- -ffast-math
- CXXOPTIMIZE
- Enables a range of optimizations that provide faster, though sometimes less precise, mathematical operations that may not conform to the IEEE-754 specifications. When this option is specified, the __STDC_IEC_559__ macro is ignored even if set by the system headers.
- -march=znver3
- CXXOPTIMIZE
- Specify that Clang should generate code for a specific processor family member and later. For example, if you specify -march=znver1, the compiler is allowed to generate instructions that are valid on AMD Zen processors, but which may not exist on earlier products.
- -fveclib=AMDLIBM
- CXXOPTIMIZE
- Use the given vector functions library.
- -mllvm -enable-partial-unswitch
- CXXOPTIMIZE
- This optimization does partial unswitching of loops where some part of the unswitched control flow remains in the loop.
- -mllvm -unroll-threshold=100
- CXXOPTIMIZE
- Sets the limit at which loops will be unrolled. For example, if unroll threshold is set to 100 then only loops with 100 or fewer instructions will be unrolled.
- -finline-aggressive
- CXXOPTIMIZE
- Sets the compiler's inlining heuristics to an aggressive level by increasing the inline thresholds.
- -flv-function-specialization
- CXXOPTIMIZE
- This option enables an optimization that generates and calls specialized function versions when the loops inside function are vectorizable and the arguments are not aliased with each other. This optimization helps in function inlining and vectorization.
- -mllvm -loop-unswitch-threshold=200000
- CXXOPTIMIZE
- Sets the limit at which loops will be unswitched. For example, if unswitch threshold is set to 100 then only loops with 100 or fewer instructions will be unswtched.
- -mllvm -reroll-loops
- CXXOPTIMIZE
- Run the loop rerolling pass.
- -mllvm -aggressive-loop-unswitch
- CXXOPTIMIZE
- This option enables aggressive loop unswitching heuristic (including -enable-partial-unswitch) based on the usage of the branch conditional values. Loop unswitching leads to code-bloat. Code-bloat can be minimized if the hoisted condition is executed more often. This heuristic prioritizes the conditions based on the number of times they are used within the loop. The heuristic can be controlled with the following options:
  - -unswitch-identical-branches-min-count=<n>
    Enables unswitching of a loop with respect to a branch conditional value (B), where B appears in at least <n> compares in the loop. This option is enabled with -aggressive-loop-unswitch. Default value is 3.
    
    Usage: -mllvm -aggressive-loop-unswitch -mllvm -unswitch-identical-branches-min-count=<n> where n is a positive integer and lower value of <n> facilitates more unswitching
  - -unswitch-identical-branches-max-count=<n>
    Enables unswitching of a loop with respect to a branch conditional value (B), where B appears in at most <n> compares in the loop. This option is enabled with -aggressive-loop-unswitch. Default value is 6.
    
    Usage: -mllvm -aggressive-loop-unswitch -mllvm -unswitch-identical-branches-max-count=<n> where n is a positive integer and higher value of <n> facilitates more unswitching
  Note: These options may facilitate more unswitching in some of the workloads. Since loop-unswitching inherently leads to code bloat, facilitating more unswitching may significantly increase the code size and hence may also lead to longer compilation times.
- Includes:
  - -Wl,-mllvm -Wl,-enable-partial-unswitch
- -mllvm -extra-vectorizer-passes
- CXXOPTIMIZE
- Run cleanup optimization passes after vectorization.
- -mllvm -reduce-array-computations=3
- CXXOPTIMIZE
- This option eliminates the array computations based on their usage. The computations on unused array elements and computations on zero valued array elements are eliminated with this optimization. -flto as whole program analysis is required to perform this optimization.
  
  Possible values:
  - 1: Eliminates the computations on unused array elements
  - 2: Eliminates the computations on zero valued array elements
  - 3: Eliminates the computations on unused and zero valued array elements
- -mllvm -global-vectorize-slp=true
- CXXOPTIMIZE
- This option enables an optimization that does the slp vectorization across basic blocks. The SLP vectorizer vectorizes instructions within basic blocks. The global slp vectorizer analyzes instructions across basic blocks and vectorizes them.
- -mllvm -convert-pow-exp-to-int=false
- CXXOPTIMIZE
- Converts the call to floating point exponent version of pow to its integer exponent version if the floating-point exponent can be converted to integer. This option is set to true by default.
- -z muldefs
- LDOPTIMIZE
- Instructs the linker to use the first definition encountered for a symbol, and ignore all others.
- -mllvm -do-block-reorder=aggressive
- EXTRA_CXXFLAGS
- Block Reordering
  
  Possible values:
  - none: No block reorder
  - simple: Simple block reorder
  - aggressive: Aggressive block reorder
- -fvirtual-function-elimination
- EXTRA_CXXFLAGS
- Enables dead virtual function elimination optimization. Requires -flto=full.
- -fvisibility=hidden
- EXTRA_CXXFLAGS
- Set the default symbol visibility for all global declarations.
- -lamdlibm
- EXTRA_LIBS
- Instructs the compiler to link with AMD-supported optimized math library.
- -ljemalloc
- EXTRA_LIBS
- Use the jemalloc library, which is a general purpose malloc(3) implementation that emphasizes fragmentation avoidance and scalable concurrency support.
- -lflang
- EXTRA_LIBS
- Instructs the compiler to link with flang Fortran runtime libraries.
- -lflangrti
- EXTRA_LIBS
- Instructs the compiler to link with flang Fortran runtime libraries.

Fortran benchmarks

- -m64
- FC, LD
- Generate code for a 64-bit environment. The 64-bit environment sets int to 32 bits and long and pointer to 64 bits and generates code for AMD's x86-64 architecture. The compiler generates AMD64, INTEL64, x86-64 64-bit ABI. The default on a 32-bit host is 32-bit ABI. The default on a 64-bit host is 64-bit ABI if the target platform specified is 64-bit, otherwise the default is 32-bit.
- -Wl,-mllvm -Wl,-inline-recursion=4
- LDFFLAGS
- Enables inlining for recursive functions based on heuristics, with level 4 being most aggressive. Higher levels may lead to code bloat due to expansion of recursive functions at call sites.
  
  Levels:
  - 0 [DEFAULT]: Disables inlining for recursive functions.
  - 1: Enables inlining for recursive functions using heuristics with inline depth 1.
  - 2: Same as level 1 but with more aggressive heuristics.
  - 3: Enables inlining for all recursive functions with inline depth 1.
  - 4: Enables inlining for all recursive function with inline depth 10.
- -Wl,-mllvm -Wl,-lsr-in-nested-loop
- LDFFLAGS
- Enables loop strength reduction for nested loop structures. By default, the compiler will do loop strength reduction only for the innermost loop.
- -Wl,-mllvm -Wl,-enable-iv-split
- LDFFLAGS
- Enables splitting of long live ranges of loop induction variables which span loop boundaries. This helps reduce register pressure and can help avoid needless spills to memory and reloads from memory.
- -flto
- EXTRA_LDFLAGS, FOPTIMIZE
- Generate output files in LLVM formats suitable for link time optimization. When used with -S this generates LLVM intermediate language assembly files, otherwise this generates LLVM bitcode format object files (which may be passed to the linker depending on the stage selection options).
- -Wl,-mllvm -Wl,-region-vectorize
- EXTRA_LDFLAGS
- This flag enables vectorization of loops with complex control flow that can not be vectorized by loop and slp vectorizers.
- -Wl,-mllvm -Wl,-function-specialize
- EXTRA_LDFLAGS
- This option enables an optimization that generates and calls specialized function versions when they are called with constant arguments. This optimization helps in function inlining.
- -Wl,-mllvm -Wl,-align-all-nofallthru-blocks=6
- EXTRA_LDFLAGS
- Force the alignment of all blocks that have no fall-through predecessors (i.e. don't add nops that are executed). In log2 format (e.g 4 means align on 16B boundaries).
- -Wl,-mllvm -Wl,-reduce-array-computations=3
- EXTRA_LDFLAGS
- This option eliminates the array computations based on their usage. The computations on unused array elements and computations on zero valued array elements are eliminated with this optimization. -flto as whole program analysis is required to perform this optimization.
  
  Possible values:
  - 1: Eliminates the computations on unused array elements
  - 2: Eliminates the computations on zero valued array elements
  - 3: Eliminates the computations on unused and zero valued array elements
- -O3
- FOPTIMIZE
- Like -O2, except that it enables optimizations that take longer to perform or that may generate larger code (in an attempt to make the program run faster).
  
  If multiple "O" options are used, with or without level numbers, the last such option is the one that is effective.
- Includes:
  - -O2
    - -O1
- -ffast-math
- FOPTIMIZE
- Enables a range of optimizations that provide faster, though sometimes less precise, mathematical operations that may not conform to the IEEE-754 specifications. When this option is specified, the __STDC_IEC_559__ macro is ignored even if set by the system headers.
- -march=znver3
- FOPTIMIZE
- Specify that Clang should generate code for a specific processor family member and later. For example, if you specify -march=znver1, the compiler is allowed to generate instructions that are valid on AMD Zen processors, but which may not exist on earlier products.
- -fveclib=AMDLIBM
- FOPTIMIZE
- Use the given vector functions library.
- -z muldefs
- LDOPTIMIZE
- Instructs the linker to use the first definition encountered for a symbol, and ignore all others.
- -mllvm -unroll-aggressive
- EXTRA_FFLAGS
- Enables aggressive heuristics to get loop unrolling.
- -mllvm -unroll-threshold=500
- EXTRA_FFLAGS
- Sets the limit at which loops will be unrolled. For example, if unroll threshold is set to 100 then only loops with 100 or fewer instructions will be unrolled.
- -lamdlibm
- EXTRA_LIBS
- Instructs the compiler to link with AMD-supported optimized math library.
- -ljemalloc
- EXTRA_LIBS
- Use the jemalloc library, which is a general purpose malloc(3) implementation that emphasizes fragmentation avoidance and scalable concurrency support.
- -lflang
- EXTRA_LIBS
- Instructs the compiler to link with flang Fortran runtime libraries.
- -lflangrti
- EXTRA_LIBS
- Instructs the compiler to link with flang Fortran runtime libraries.

Peak Optimization Flags

C benchmarks

500.perlbench_r

- -m64
- CC, LD
- Generate code for a 64-bit environment. The 64-bit environment sets int to 32 bits and long and pointer to 64 bits and generates code for AMD's x86-64 architecture. The compiler generates AMD64, INTEL64, x86-64 64-bit ABI. The default on a 32-bit host is 32-bit ABI. The default on a 64-bit host is 64-bit ABI if the target platform specified is 64-bit, otherwise the default is 32-bit.
- -Wl,-allow-multiple-definition
- LDCFLAGS
- Do not generate an error when linking multiple symbols of the same name.
- -Wl,-mllvm -Wl,-enable-licm-vrp
- LDCFLAGS
- Enables estimation of the virtual register pressure before performing loop invariant code motion. This estimation is used to decide the invariants that will be hoisted during loop invariant code motion.
- -flto
- COPTIMIZE, EXTRA_LDFLAGS
- Generate output files in LLVM formats suitable for link time optimization. When used with -S this generates LLVM intermediate language assembly files, otherwise this generates LLVM bitcode format object files (which may be passed to the linker depending on the stage selection options).
- -Wl,-mllvm -Wl,-function-specialize
- EXTRA_LDFLAGS
- This option enables an optimization that generates and calls specialized function versions when they are called with constant arguments. This optimization helps in function inlining.
- -Wl,-mllvm -Wl,-align-all-nofallthru-blocks=6
- EXTRA_LDFLAGS
- Force the alignment of all blocks that have no fall-through predecessors (i.e. don't add nops that are executed). In log2 format (e.g 4 means align on 16B boundaries).
- -Wl,-mllvm -Wl,-reduce-array-computations=3
- EXTRA_LDFLAGS
- This option eliminates the array computations based on their usage. The computations on unused array elements and computations on zero valued array elements are eliminated with this optimization. -flto as whole program analysis is required to perform this optimization.
  
  Possible values:
  - 1: Eliminates the computations on unused array elements
  - 2: Eliminates the computations on zero valued array elements
  - 3: Eliminates the computations on unused and zero valued array elements
- -fprofile-instr-generate
- PASS1_CFLAGS, PASS1_LDFLAGS
- Turns on LLVM's instrumenation based profiling.
- -fprofile-instr-use
- PASS2_CFLAGS, PASS2_LDFLAGS
- Uses the profiling files generated from a program compiled with -fprofile-instr-generate to guide optimization decisions.
- -Ofast
- COPTIMIZE
- Enables all the optimizations from -O3 along with other aggressive optimizations that may violate strict compliance with language standards. Refer to the AOCC options document for the language you're using for more detailed documentation of optimizations enabled under -Ofast.
- Includes:
  - -O3
    - -O2
      
      -O1
- -march=znver3
- COPTIMIZE
- Specify that Clang should generate code for a specific processor family member and later. For example, if you specify -march=znver1, the compiler is allowed to generate instructions that are valid on AMD Zen processors, but which may not exist on earlier products.
- -fveclib=AMDLIBM
- COPTIMIZE
- Use the given vector functions library.
- -fstruct-layout=7
- COPTIMIZE
- Analyzes the whole program to determine if the structures in the code can be peeled and if pointer or integer fields in the structure can be compressed. If feasible, this optimization transforms the code to enable these improvements. This transformation is likely to improve cache utilization and memory bandwidth. This, in turn, is expected to improve the scalability of programs executed on multiple cores.
  
  This is effective only under -flto as whole program analysis is required to perform this optimization. You can choose different levels of aggressiveness with which this optimization can be applied to your application with 1 being the least aggressive and 7 being the most aggressive level.
  
  Possible values:
  - fstruct-layout=1: enables structure peeling.
  - fstruct-layout=2: enables structure peeling and selectively compresses self-referential pointers in these structures to 32-bit pointers wherever safe.
  - fstruct-layout=3: enables structure peeling and selectively compresses self-referential pointers in these structures to 16-bit pointers wherever safe.
  - fstruct-layout=4: enables structure peeling, pointer compression as in level 2 and further enables compression of structure fields which are of integer type. This is performed under a strict safety check.
  - fstruct-layout=5: enables structure peeling, pointer compression as in level 3 and further enables compression of structure fields which are of integer type. This is performed under a strict safety check.
  - fstruct-layout=6: enables structure peeling, pointer compression as in level 2 and further enables compression of structure fields which are of type 64-bit 'signed int' or 'unsigned Int'. The user needs to ensure that the values assigned to 64-bit 'signed int' fields are in range -(2^31 - 1) to +(2^31 - 1) and 64-bit 'unsigned int' fields are in range 0 to +(2^31 - 1), otherwise, incorrect results may be obtained. This compression is performed without considering any safety analysis and so the user needs to ensure the safety based on the program compiled.
  - fstruct-layout=7: enables structure peeling, pointer compression as in level 3 and further enables compression of structure fields which are of type 64-bit 'signed int' or 'unsigned Int'. The user needs to ensure that the values assigned to 64-bit 'signed int' fields are in range -(2^31 - 1) to +(2^31 - 1) and 64-bit 'unsigned int' field are in range 0 to +(2^31 - 1) , otherwise, incorrect results may be obtained. This compression is performed without considering any safety analysis and so the user needs to ensure the safety based on the program compiled.
  Note:
  fstruct-layout=4 and fstruct-layout=5 are derived from fstruct-layout=2 and fstruct-layout=3 respectively with the added feature of safe compression of integer fields in structures. Going from fstruct-layout=4 to fstruct-layout=5 may result in higher performance if the pointer values are such that the pointers can be compressed to 16-bits.
  
  fstruct-layout=6 and fstruct-layout=7 are derived from fstruct-layout=2 and fstruct-layout=3 respectively with the added feature of compression of integer fields in structures. These are similar to fstruct-layout=4 and fstruct-layout=5, but here, the integer fields of the structures are always compressed from 64-bits to 32-bits without any safety guarantee.
- -mllvm -unroll-threshold=50
- COPTIMIZE
- Sets the limit at which loops will be unrolled. For example, if unroll threshold is set to 100 then only loops with 100 or fewer instructions will be unrolled.
- -fremap-arrays
- COPTIMIZE
- This option enables an optimization that transforms the data layout of a single dimensional array to provide better cache locality by analysing the access patterns.
- -flv-function-specialization
- COPTIMIZE
- This option enables an optimization that generates and calls specialized function versions when the loops inside function are vectorizable and the arguments are not aliased with each other. This optimization helps in function inlining and vectorization.
- -mllvm -inline-threshold=1000
- COPTIMIZE
- Sets the compiler's inlining threshold level to the value passed as the argument. The inline threshold is used in the inliner heuristics to decide which functions should be inlined.
- -mllvm -enable-gvn-hoist
- COPTIMIZE
- This option enables the GVN hoist pass, which is used to hoist computations from branches.
- -mllvm -global-vectorize-slp=false
- COPTIMIZE
- This option enables an optimization that does the slp vectorization across basic blocks. The SLP vectorizer vectorizes instructions within basic blocks. The global slp vectorizer analyzes instructions across basic blocks and vectorizes them.
- -mllvm -function-specialize
- COPTIMIZE
- This option enables an optimization that generates and calls specialized function versions when they are called with constant arguments. This optimization helps in function inlining.
- -mllvm -enable-licm-vrp
- COPTIMIZE
- Enables estimation of the virtual register pressure before performing loop invariant code motion. This estimation is used to decide the invariants that will be hoisted during loop invariant code motion.
- -mllvm -reduce-array-computations=3
- COPTIMIZE
- This option eliminates the array computations based on their usage. The computations on unused array elements and computations on zero valued array elements are eliminated with this optimization. -flto as whole program analysis is required to perform this optimization.
  
  Possible values:
  - 1: Eliminates the computations on unused array elements
  - 2: Eliminates the computations on zero valued array elements
  - 3: Eliminates the computations on unused and zero valued array elements
- -lamdlibm
- EXTRA_LIBS
- Instructs the compiler to link with AMD-supported optimized math library.
- -ljemalloc
- EXTRA_LIBS
- Use the jemalloc library, which is a general purpose malloc(3) implementation that emphasizes fragmentation avoidance and scalable concurrency support.

502.gcc_r

- -m32
- CC, LD
- Generate code for a 32-bit environment. The 32-bit environment sets int, long and pointer to 32 bits and generates code that runs on any i386 system. The compiler generates x86 or IA32 32-bit ABI. The default on a 32-bit host is 32-bit ABI. The default on a 64-bit host is 64-bit ABI if the target platform specified is 64-bit, otherwise the default is 32-bit.
- -Wl,-allow-multiple-definition
- LDCFLAGS
- Do not generate an error when linking multiple symbols of the same name.
- -Wl,-mllvm -Wl,-enable-licm-vrp
- LDCFLAGS
- Enables estimation of the virtual register pressure before performing loop invariant code motion. This estimation is used to decide the invariants that will be hoisted during loop invariant code motion.
- -flto
- COPTIMIZE, EXTRA_LDFLAGS
- Generate output files in LLVM formats suitable for link time optimization. When used with -S this generates LLVM intermediate language assembly files, otherwise this generates LLVM bitcode format object files (which may be passed to the linker depending on the stage selection options).
- -Wl,-mllvm -Wl,-function-specialize
- EXTRA_LDFLAGS
- This option enables an optimization that generates and calls specialized function versions when they are called with constant arguments. This optimization helps in function inlining.
- -Ofast
- COPTIMIZE
- Enables all the optimizations from -O3 along with other aggressive optimizations that may violate strict compliance with language standards. Refer to the AOCC options document for the language you're using for more detailed documentation of optimizations enabled under -Ofast.
- Includes:
  - -O3
    - -O2
      
      -O1
- -march=znver3
- COPTIMIZE
- Specify that Clang should generate code for a specific processor family member and later. For example, if you specify -march=znver1, the compiler is allowed to generate instructions that are valid on AMD Zen processors, but which may not exist on earlier products.
- -fveclib=AMDLIBM
- COPTIMIZE
- Use the given vector functions library.
- -fstruct-layout=7
- COPTIMIZE
- Analyzes the whole program to determine if the structures in the code can be peeled and if pointer or integer fields in the structure can be compressed. If feasible, this optimization transforms the code to enable these improvements. This transformation is likely to improve cache utilization and memory bandwidth. This, in turn, is expected to improve the scalability of programs executed on multiple cores.
  
  This is effective only under -flto as whole program analysis is required to perform this optimization. You can choose different levels of aggressiveness with which this optimization can be applied to your application with 1 being the least aggressive and 7 being the most aggressive level.
  
  Possible values:
  - fstruct-layout=1: enables structure peeling.
  - fstruct-layout=2: enables structure peeling and selectively compresses self-referential pointers in these structures to 32-bit pointers wherever safe.
  - fstruct-layout=3: enables structure peeling and selectively compresses self-referential pointers in these structures to 16-bit pointers wherever safe.
  - fstruct-layout=4: enables structure peeling, pointer compression as in level 2 and further enables compression of structure fields which are of integer type. This is performed under a strict safety check.
  - fstruct-layout=5: enables structure peeling, pointer compression as in level 3 and further enables compression of structure fields which are of integer type. This is performed under a strict safety check.
  - fstruct-layout=6: enables structure peeling, pointer compression as in level 2 and further enables compression of structure fields which are of type 64-bit 'signed int' or 'unsigned Int'. The user needs to ensure that the values assigned to 64-bit 'signed int' fields are in range -(2^31 - 1) to +(2^31 - 1) and 64-bit 'unsigned int' fields are in range 0 to +(2^31 - 1), otherwise, incorrect results may be obtained. This compression is performed without considering any safety analysis and so the user needs to ensure the safety based on the program compiled.
  - fstruct-layout=7: enables structure peeling, pointer compression as in level 3 and further enables compression of structure fields which are of type 64-bit 'signed int' or 'unsigned Int'. The user needs to ensure that the values assigned to 64-bit 'signed int' fields are in range -(2^31 - 1) to +(2^31 - 1) and 64-bit 'unsigned int' field are in range 0 to +(2^31 - 1) , otherwise, incorrect results may be obtained. This compression is performed without considering any safety analysis and so the user needs to ensure the safety based on the program compiled.
  Note:
  fstruct-layout=4 and fstruct-layout=5 are derived from fstruct-layout=2 and fstruct-layout=3 respectively with the added feature of safe compression of integer fields in structures. Going from fstruct-layout=4 to fstruct-layout=5 may result in higher performance if the pointer values are such that the pointers can be compressed to 16-bits.
  
  fstruct-layout=6 and fstruct-layout=7 are derived from fstruct-layout=2 and fstruct-layout=3 respectively with the added feature of compression of integer fields in structures. These are similar to fstruct-layout=4 and fstruct-layout=5, but here, the integer fields of the structures are always compressed from 64-bits to 32-bits without any safety guarantee.
- -mllvm -unroll-threshold=50
- COPTIMIZE
- Sets the limit at which loops will be unrolled. For example, if unroll threshold is set to 100 then only loops with 100 or fewer instructions will be unrolled.
- -fremap-arrays
- COPTIMIZE
- This option enables an optimization that transforms the data layout of a single dimensional array to provide better cache locality by analysing the access patterns.
- -flv-function-specialization
- COPTIMIZE
- This option enables an optimization that generates and calls specialized function versions when the loops inside function are vectorizable and the arguments are not aliased with each other. This optimization helps in function inlining and vectorization.
- -mllvm -inline-threshold=1000
- COPTIMIZE
- Sets the compiler's inlining threshold level to the value passed as the argument. The inline threshold is used in the inliner heuristics to decide which functions should be inlined.
- -mllvm -enable-gvn-hoist
- COPTIMIZE
- This option enables the GVN hoist pass, which is used to hoist computations from branches.
- -mllvm -global-vectorize-slp=true
- COPTIMIZE
- This option enables an optimization that does the slp vectorization across basic blocks. The SLP vectorizer vectorizes instructions within basic blocks. The global slp vectorizer analyzes instructions across basic blocks and vectorizes them.
- -mllvm -function-specialize
- COPTIMIZE
- This option enables an optimization that generates and calls specialized function versions when they are called with constant arguments. This optimization helps in function inlining.
- -mllvm -enable-licm-vrp
- COPTIMIZE
- Enables estimation of the virtual register pressure before performing loop invariant code motion. This estimation is used to decide the invariants that will be hoisted during loop invariant code motion.
- -mllvm -reduce-array-computations=3
- COPTIMIZE
- This option eliminates the array computations based on their usage. The computations on unused array elements and computations on zero valued array elements are eliminated with this optimization. -flto as whole program analysis is required to perform this optimization.
  
  Possible values:
  - 1: Eliminates the computations on unused array elements
  - 2: Eliminates the computations on zero valued array elements
  - 3: Eliminates the computations on unused and zero valued array elements
- -fgnu89-inline
- EXTRA_COPTIMIZE
- In the 502/602.gcc benchmark description, "multiple definitions of symbols" is listed under the "Known Portability Issues" section, and this option is one of the suggested workarounds. This option causes Clang to revert to the same inlining behavior that GCC does when in pre-C99 mode.
- -ljemalloc
- EXTRA_LIBS
- Use the jemalloc library, which is a general purpose malloc(3) implementation that emphasizes fragmentation avoidance and scalable concurrency support.

C++ benchmarks

523.xalancbmk_r

- -m32
- CXX, LD
- Generate code for a 32-bit environment. The 32-bit environment sets int, long and pointer to 32 bits and generates code that runs on any i386 system. The compiler generates x86 or IA32 32-bit ABI. The default on a 32-bit host is 32-bit ABI. The default on a 64-bit host is 64-bit ABI if the target platform specified is 64-bit, otherwise the default is 32-bit.
- -Wl,-mllvm -Wl,-do-block-reorder=aggressive
- LDCXXFLAGS
- Block Reordering
  
  Possible values:
  - none: No block reorder
  - simple: Simple block reorder
  - aggressive: Aggressive block reorder
- -flto
- CXXOPTIMIZE, EXTRA_LDFLAGS
- Generate output files in LLVM formats suitable for link time optimization. When used with -S this generates LLVM intermediate language assembly files, otherwise this generates LLVM bitcode format object files (which may be passed to the linker depending on the stage selection options).
- -Wl,-mllvm -Wl,-function-specialize
- EXTRA_LDFLAGS
- This option enables an optimization that generates and calls specialized function versions when they are called with constant arguments. This optimization helps in function inlining.
- -Wl,-mllvm -Wl,-align-all-nofallthru-blocks=6
- EXTRA_LDFLAGS
- Force the alignment of all blocks that have no fall-through predecessors (i.e. don't add nops that are executed). In log2 format (e.g 4 means align on 16B boundaries).
- -Wl,-mllvm -Wl,-reduce-array-computations=3
- EXTRA_LDFLAGS
- This option eliminates the array computations based on their usage. The computations on unused array elements and computations on zero valued array elements are eliminated with this optimization. -flto as whole program analysis is required to perform this optimization.
  
  Possible values:
  - 1: Eliminates the computations on unused array elements
  - 2: Eliminates the computations on zero valued array elements
  - 3: Eliminates the computations on unused and zero valued array elements
- -Ofast
- CXXOPTIMIZE
- Enables all the optimizations from -O3 along with other aggressive optimizations that may violate strict compliance with language standards. Refer to the AOCC options document for the language you're using for more detailed documentation of optimizations enabled under -Ofast.
- Includes:
  - -O3
    - -O2
      
      -O1
- -march=znver3
- CXXOPTIMIZE
- Specify that Clang should generate code for a specific processor family member and later. For example, if you specify -march=znver1, the compiler is allowed to generate instructions that are valid on AMD Zen processors, but which may not exist on earlier products.
- -fveclib=AMDLIBM
- CXXOPTIMIZE
- Use the given vector functions library.
- -finline-aggressive
- CXXOPTIMIZE
- Sets the compiler's inlining heuristics to an aggressive level by increasing the inline thresholds.
- -mllvm -unroll-threshold=100
- CXXOPTIMIZE
- Sets the limit at which loops will be unrolled. For example, if unroll threshold is set to 100 then only loops with 100 or fewer instructions will be unrolled.
- -flv-function-specialization
- CXXOPTIMIZE
- This option enables an optimization that generates and calls specialized function versions when the loops inside function are vectorizable and the arguments are not aliased with each other. This optimization helps in function inlining and vectorization.
- -mllvm -enable-licm-vrp
- CXXOPTIMIZE
- Enables estimation of the virtual register pressure before performing loop invariant code motion. This estimation is used to decide the invariants that will be hoisted during loop invariant code motion.
- -mllvm -reroll-loops
- CXXOPTIMIZE
- Run the loop rerolling pass.
- -mllvm -aggressive-loop-unswitch
- CXXOPTIMIZE
- This option enables aggressive loop unswitching heuristic (including -enable-partial-unswitch) based on the usage of the branch conditional values. Loop unswitching leads to code-bloat. Code-bloat can be minimized if the hoisted condition is executed more often. This heuristic prioritizes the conditions based on the number of times they are used within the loop. The heuristic can be controlled with the following options:
  - -unswitch-identical-branches-min-count=<n>
    Enables unswitching of a loop with respect to a branch conditional value (B), where B appears in at least <n> compares in the loop. This option is enabled with -aggressive-loop-unswitch. Default value is 3.
    
    Usage: -mllvm -aggressive-loop-unswitch -mllvm -unswitch-identical-branches-min-count=<n> where n is a positive integer and lower value of <n> facilitates more unswitching
  - -unswitch-identical-branches-max-count=<n>
    Enables unswitching of a loop with respect to a branch conditional value (B), where B appears in at most <n> compares in the loop. This option is enabled with -aggressive-loop-unswitch. Default value is 6.
    
    Usage: -mllvm -aggressive-loop-unswitch -mllvm -unswitch-identical-branches-max-count=<n> where n is a positive integer and higher value of <n> facilitates more unswitching
  Note: These options may facilitate more unswitching in some of the workloads. Since loop-unswitching inherently leads to code bloat, facilitating more unswitching may significantly increase the code size and hence may also lead to longer compilation times.
- Includes:
  - -Wl,-mllvm -Wl,-enable-partial-unswitch
- -mllvm -reduce-array-computations=3
- CXXOPTIMIZE
- This option eliminates the array computations based on their usage. The computations on unused array elements and computations on zero valued array elements are eliminated with this optimization. -flto as whole program analysis is required to perform this optimization.
  
  Possible values:
  - 1: Eliminates the computations on unused array elements
  - 2: Eliminates the computations on zero valued array elements
  - 3: Eliminates the computations on unused and zero valued array elements
- -mllvm -global-vectorize-slp=true
- CXXOPTIMIZE
- This option enables an optimization that does the slp vectorization across basic blocks. The SLP vectorizer vectorizes instructions within basic blocks. The global slp vectorizer analyzes instructions across basic blocks and vectorizes them.
- -mllvm -do-block-reorder=aggressive
- EXTRA_CXXFLAGS
- Block Reordering
  
  Possible values:
  - none: No block reorder
  - simple: Simple block reorder
  - aggressive: Aggressive block reorder
- -fvirtual-function-elimination
- EXTRA_CXXFLAGS
- Enables dead virtual function elimination optimization. Requires -flto=full.
- -fvisibility=hidden
- EXTRA_CXXFLAGS
- Set the default symbol visibility for all global declarations.
- -ljemalloc
- EXTRA_LIBS
- Use the jemalloc library, which is a general purpose malloc(3) implementation that emphasizes fragmentation avoidance and scalable concurrency support.

531.deepsjeng_r

- -m64
- CXX, LD
- Generate code for a 64-bit environment. The 64-bit environment sets int to 32 bits and long and pointer to 64 bits and generates code for AMD's x86-64 architecture. The compiler generates AMD64, INTEL64, x86-64 64-bit ABI. The default on a 32-bit host is 32-bit ABI. The default on a 64-bit host is 64-bit ABI if the target platform specified is 64-bit, otherwise the default is 32-bit.
- -std=c++98
- CXX
- Selects the C++ language dialect.
- -Wl,-mllvm -Wl,-do-block-reorder=aggressive
- LDCXXFLAGS
- Block Reordering
  
  Possible values:
  - none: No block reorder
  - simple: Simple block reorder
  - aggressive: Aggressive block reorder
- -flto
- CXXOPTIMIZE, EXTRA_LDFLAGS
- Generate output files in LLVM formats suitable for link time optimization. When used with -S this generates LLVM intermediate language assembly files, otherwise this generates LLVM bitcode format object files (which may be passed to the linker depending on the stage selection options).
- -Wl,-mllvm -Wl,-function-specialize
- EXTRA_LDFLAGS
- This option enables an optimization that generates and calls specialized function versions when they are called with constant arguments. This optimization helps in function inlining.
- -Wl,-mllvm -Wl,-align-all-nofallthru-blocks=6
- EXTRA_LDFLAGS
- Force the alignment of all blocks that have no fall-through predecessors (i.e. don't add nops that are executed). In log2 format (e.g 4 means align on 16B boundaries).
- -Wl,-mllvm -Wl,-reduce-array-computations=3
- EXTRA_LDFLAGS
- This option eliminates the array computations based on their usage. The computations on unused array elements and computations on zero valued array elements are eliminated with this optimization. -flto as whole program analysis is required to perform this optimization.
  
  Possible values:
  - 1: Eliminates the computations on unused array elements
  - 2: Eliminates the computations on zero valued array elements
  - 3: Eliminates the computations on unused and zero valued array elements
- -Ofast
- CXXOPTIMIZE
- Enables all the optimizations from -O3 along with other aggressive optimizations that may violate strict compliance with language standards. Refer to the AOCC options document for the language you're using for more detailed documentation of optimizations enabled under -Ofast.
- Includes:
  - -O3
    - -O2
      
      -O1
- -march=znver3
- CXXOPTIMIZE
- Specify that Clang should generate code for a specific processor family member and later. For example, if you specify -march=znver1, the compiler is allowed to generate instructions that are valid on AMD Zen processors, but which may not exist on earlier products.
- -fveclib=AMDLIBM
- CXXOPTIMIZE
- Use the given vector functions library.
- -finline-aggressive
- CXXOPTIMIZE
- Sets the compiler's inlining heuristics to an aggressive level by increasing the inline thresholds.
- -mllvm -unroll-threshold=100
- CXXOPTIMIZE
- Sets the limit at which loops will be unrolled. For example, if unroll threshold is set to 100 then only loops with 100 or fewer instructions will be unrolled.
- -flv-function-specialization
- CXXOPTIMIZE
- This option enables an optimization that generates and calls specialized function versions when the loops inside function are vectorizable and the arguments are not aliased with each other. This optimization helps in function inlining and vectorization.
- -mllvm -enable-licm-vrp
- CXXOPTIMIZE
- Enables estimation of the virtual register pressure before performing loop invariant code motion. This estimation is used to decide the invariants that will be hoisted during loop invariant code motion.
- -mllvm -reroll-loops
- CXXOPTIMIZE
- Run the loop rerolling pass.
- -mllvm -aggressive-loop-unswitch
- CXXOPTIMIZE
- This option enables aggressive loop unswitching heuristic (including -enable-partial-unswitch) based on the usage of the branch conditional values. Loop unswitching leads to code-bloat. Code-bloat can be minimized if the hoisted condition is executed more often. This heuristic prioritizes the conditions based on the number of times they are used within the loop. The heuristic can be controlled with the following options:
  - -unswitch-identical-branches-min-count=<n>
    Enables unswitching of a loop with respect to a branch conditional value (B), where B appears in at least <n> compares in the loop. This option is enabled with -aggressive-loop-unswitch. Default value is 3.
    
    Usage: -mllvm -aggressive-loop-unswitch -mllvm -unswitch-identical-branches-min-count=<n> where n is a positive integer and lower value of <n> facilitates more unswitching
  - -unswitch-identical-branches-max-count=<n>
    Enables unswitching of a loop with respect to a branch conditional value (B), where B appears in at most <n> compares in the loop. This option is enabled with -aggressive-loop-unswitch. Default value is 6.
    
    Usage: -mllvm -aggressive-loop-unswitch -mllvm -unswitch-identical-branches-max-count=<n> where n is a positive integer and higher value of <n> facilitates more unswitching
  Note: These options may facilitate more unswitching in some of the workloads. Since loop-unswitching inherently leads to code bloat, facilitating more unswitching may significantly increase the code size and hence may also lead to longer compilation times.
- Includes:
  - -Wl,-mllvm -Wl,-enable-partial-unswitch
- -mllvm -reduce-array-computations=3
- CXXOPTIMIZE
- This option eliminates the array computations based on their usage. The computations on unused array elements and computations on zero valued array elements are eliminated with this optimization. -flto as whole program analysis is required to perform this optimization.
  
  Possible values:
  - 1: Eliminates the computations on unused array elements
  - 2: Eliminates the computations on zero valued array elements
  - 3: Eliminates the computations on unused and zero valued array elements
- -mllvm -global-vectorize-slp=true
- CXXOPTIMIZE
- This option enables an optimization that does the slp vectorization across basic blocks. The SLP vectorizer vectorizes instructions within basic blocks. The global slp vectorizer analyzes instructions across basic blocks and vectorizes them.
- -mllvm -do-block-reorder=aggressive
- EXTRA_CXXFLAGS
- Block Reordering
  
  Possible values:
  - none: No block reorder
  - simple: Simple block reorder
  - aggressive: Aggressive block reorder
- -fvirtual-function-elimination
- EXTRA_CXXFLAGS
- Enables dead virtual function elimination optimization. Requires -flto=full.
- -fvisibility=hidden
- EXTRA_CXXFLAGS
- Set the default symbol visibility for all global declarations.
- -lamdlibm
- EXTRA_LIBS
- Instructs the compiler to link with AMD-supported optimized math library.
- -ljemalloc
- EXTRA_LIBS
- Use the jemalloc library, which is a general purpose malloc(3) implementation that emphasizes fragmentation avoidance and scalable concurrency support.

Fortran benchmarks

- -m64
- FC, LD
- Generate code for a 64-bit environment. The 64-bit environment sets int to 32 bits and long and pointer to 64 bits and generates code for AMD's x86-64 architecture. The compiler generates AMD64, INTEL64, x86-64 64-bit ABI. The default on a 32-bit host is 32-bit ABI. The default on a 64-bit host is 64-bit ABI if the target platform specified is 64-bit, otherwise the default is 32-bit.
- -Wl,-mllvm -Wl,-inline-recursion=4
- LDFFLAGS
- Enables inlining for recursive functions based on heuristics, with level 4 being most aggressive. Higher levels may lead to code bloat due to expansion of recursive functions at call sites.
  
  Levels:
  - 0 [DEFAULT]: Disables inlining for recursive functions.
  - 1: Enables inlining for recursive functions using heuristics with inline depth 1.
  - 2: Same as level 1 but with more aggressive heuristics.
  - 3: Enables inlining for all recursive functions with inline depth 1.
  - 4: Enables inlining for all recursive function with inline depth 10.
- -Wl,-mllvm -Wl,-lsr-in-nested-loop
- LDFFLAGS
- Enables loop strength reduction for nested loop structures. By default, the compiler will do loop strength reduction only for the innermost loop.
- -Wl,-mllvm -Wl,-enable-iv-split
- LDFFLAGS
- Enables splitting of long live ranges of loop induction variables which span loop boundaries. This helps reduce register pressure and can help avoid needless spills to memory and reloads from memory.
- -flto
- EXTRA_LDFLAGS, FOPTIMIZE
- Generate output files in LLVM formats suitable for link time optimization. When used with -S this generates LLVM intermediate language assembly files, otherwise this generates LLVM bitcode format object files (which may be passed to the linker depending on the stage selection options).
- -Wl,-mllvm -Wl,-function-specialize
- EXTRA_LDFLAGS
- This option enables an optimization that generates and calls specialized function versions when they are called with constant arguments. This optimization helps in function inlining.
- -Wl,-mllvm -Wl,-align-all-nofallthru-blocks=6
- EXTRA_LDFLAGS
- Force the alignment of all blocks that have no fall-through predecessors (i.e. don't add nops that are executed). In log2 format (e.g 4 means align on 16B boundaries).
- -Wl,-mllvm -Wl,-reduce-array-computations=3
- EXTRA_LDFLAGS
- This option eliminates the array computations based on their usage. The computations on unused array elements and computations on zero valued array elements are eliminated with this optimization. -flto as whole program analysis is required to perform this optimization.
  
  Possible values:
  - 1: Eliminates the computations on unused array elements
  - 2: Eliminates the computations on zero valued array elements
  - 3: Eliminates the computations on unused and zero valued array elements
- -O3
- FOPTIMIZE
- Like -O2, except that it enables optimizations that take longer to perform or that may generate larger code (in an attempt to make the program run faster).
  
  If multiple "O" options are used, with or without level numbers, the last such option is the one that is effective.
- Includes:
  - -O2
    - -O1
- -ffast-math
- FOPTIMIZE
- Enables a range of optimizations that provide faster, though sometimes less precise, mathematical operations that may not conform to the IEEE-754 specifications. When this option is specified, the __STDC_IEC_559__ macro is ignored even if set by the system headers.
- -march=znver3
- FOPTIMIZE
- Specify that Clang should generate code for a specific processor family member and later. For example, if you specify -march=znver1, the compiler is allowed to generate instructions that are valid on AMD Zen processors, but which may not exist on earlier products.
- -fveclib=AMDLIBM
- FOPTIMIZE
- Use the given vector functions library.
- -mllvm -unroll-aggressive
- EXTRA_FFLAGS
- Enables aggressive heuristics to get loop unrolling.
- -mllvm -unroll-threshold=500
- EXTRA_FFLAGS
- Sets the limit at which loops will be unrolled. For example, if unroll threshold is set to 100 then only loops with 100 or fewer instructions will be unrolled.
- -lamdlibm
- EXTRA_LIBS
- Instructs the compiler to link with AMD-supported optimized math library.
- -ljemalloc
- EXTRA_LIBS
- Use the jemalloc library, which is a general purpose malloc(3) implementation that emphasizes fragmentation avoidance and scalable concurrency support.
- -lflang
- EXTRA_FLIBS
- Instructs the compiler to link with flang Fortran runtime libraries.
- -lflangrti
- EXTRA_FLIBS
- Instructs the compiler to link with flang Fortran runtime libraries.

Implicitly Included Flags

This section contains descriptions of flags that were included implicitly by other flags, but which do not have a permanent home at SPEC.

Commands and Options Used to Submit Benchmark Runs

For multi-copy runs or single copy runs on systems with multiple sockets, it is advantageous to bind a process to a particular core. Otherwise, the OS may arbitrarily move your process from one core to another. This can affect performance. To help, SPEC allows the use of a "submit" command where users can specify a utility to use to bind processes. We have found the utility 'numactl' to be the best choice.

numactl runs processes with a specific NUMA scheduling or memory placement policy. The policy is set for a command and inherited by all of its children. The numactl flag "--physcpubind" specifies which core(s) to bind the process. "-l" instructs numactl to keep a process's memory on the local node while "-m" specifies which node(s) to place a process's memory. For full details on using numactl, please refer to your Linux documentation, 'man numactl'

Note that some older versions of numactl incorrectly interpret application arguments as its own. For example, with the command "numactl --physcpubind=0 -l a.out -m a", numactl will interpret a.out's "-m" option as its own "-m" option. To work around this problem, we put the command to be run in a shell script and then run the shell script using numactl. For example: "echo 'a.out -m a' > run.sh ; numactl --physcpubind=0 bash run.sh"

Shell, Environment, and Other Software Settings

THP is an abstraction layer that automates most aspects of creating, managing, and using huge pages. THP is designed to hide much of the complexity in using huge pages from system administrators and developers, as normal huge pages must be assigned at boot time, can be difficult to manage manually, and often require significant changes to code in order to be used effectively. Most recent Linux OS releases have THP enabled by default.

Sets the stack size to n kbytes, or unlimited to allow the stack size to grow without limit.

Disables the cpu frequency scaling program in order to set the CPUs to the highest supported frequency.

An environment variable that indicates the location in the filesystem of bundled libraries to use when running the benchmark binaries.

This option can be used to select the type of process address space randomization that is used in the system, for architectures that support this feature.

An environment variable set to tune the jemalloc allocation strategy during the execution of the binaries. This environment variable setting is not needed when building the binaries on the system under test.

An environment variable used to initialize the allocated memory. Setting PGHPF_ZMEM to "Yes" has the effect of initializing all allocated memory to zero.

This environment variable is used to set the thread affinity for threads spawned by OpenMP.

This environment variable is defined as part of the OpenMP standard. Setting it to "false" prevents the OpenMP runtime from dynamically adjusting the number of threads to use for parallel execution.

This environment variable is defined as part of the OpenMP standard. Setting it to "static" causes loop iterations to be assigned to threads in round-robin fashion in the order of the thread number.

This environment variable is defined as part of the OpenMP standard and controls the size of the stack for threads created by OpenMP.

This environment variable is defined as part of the OpenMP standard and limits the maximum number of OpenMP threads that can be created.

Operating System Tuning Parameters

Firmware / BIOS / Microcode Settings

C States allow the processor to enter lower power states when idle. When set to Enabled (OS controlled) or when set to Autonomous (if Hardware controlled is supported), the processor can operate in all available Power States to save power, but my increase memory latency and frequency jitter.

This field specifies that each CCX within the processor will be declared as a NUMA Domain.

By enabling CPU memory controller to delay running the REFRESH commands, you can improve the performance for some workloads. By minimizing the delay time, it is ensured that the memory controller runs the REFRESH command at regular intervals. For Intel-based servers, this setting only affects systems configured with DIMMs which use 8Gb density DRAMs.

Most workloads will benefit from the L1 and L2 Stream Hardware prefetchers gathering data and keeping the core pipeline busy. There are however some workloads that are very random in nature and will actually obtain better overall performance by disabling one or both of the prefetchers.

Enables you to get the memory read started early on DDR bus. The Ultra Path Interconnect (UPI) Rx path will spawn the speculative memory read to Integrated Memory Controller (iMC) directly.

DLWM reduces the XGMI link width between sockets from x16 to x8 (default), when no traffic is detected on the link. This feature is optimized to trade power between core and high IO/memory bandwidth workloads.
Forced = Force link width to x16, x8, or x2.

This option configures the processor last level cache (LLC) prefetch feature as a result of the non-inclusive cache architecture. The LLC prefetcher exists on top of other prefetchers that can prefetch data into the core data cache unit (DCU) and mid-level cache (MLC). In some cases, setting this option to disabled can improve performance. Typically, setting this option to enable provides better performance. Disabled: Disables the LLC prefetcher. Enabled: Gives the core prefetcher the ability to prefetch data directly to the LLC.

In the Skylake cache scheme, mid-level cache (MLC) evictions are filled into the last level cache (LLC). If a line is evicted from the MLC to the LLC, the Skylake core can flag the evicted MLC lines as "dead". This means that the lines are not likely to be read again. This option allows dead lines to be dropped and never fill the LLC if the option is disabled. Disabled: Disabling this option can save space in the LLC by never filling dead lines into the LLC. Enabled: Opportunistically fill dead lines in LLC, if space is available.

AtoS optimization reduces remote read latencies for repeat read accesses without intervening writes.

When enabled a specific hard-fused Data Fabric (SOC) p-state is forced for optimizing workloads sensitive to latency or throughput. When disabled P-states will be automatically managed by the Application Power Management, allowing the processor to provide maximum performance while remaining within a specified power-delivery and thermal envelope.

Selecting this option allows additional cooling to the server. In case hardware is added (example, new PCIe cards), it may require additional cooling. A fan speed offset causes fan speeds to increase (by the offset % value) over baseline fan speeds calculated by the Thermal Control algorithm. Maximum — Drives fan speeds to full speed.

It controls whether BIOS will enable determinism to control performance. Performance: BIOS will enable 100% deterministic performance control. Power: BIOS will not enable deterministic performance control.

Allows selection of CPU power management methodology. Maximum Performance is typically selected for performance-centric workloads where it is acceptable to consume additional power to achieve the highest possible performance for the computing environment. This mode drives processor frequency to the maximum across all cores (although idled cores can still be frequency reduced by C-state enforcement through BIOS or OS mechanisms if enabled). This mode also offers the lowest latency of the CPU Power Management Mode options, so is always preferred for latency-sensitive environments. OS DBPM is another performance-per-watt option that relies on the operating system to dynamically control individual cores in order to save power.

Governs the BIOS memory frequency. The variables that govern maximum memory frequency include the maximum rated frequency of the DIMMs, the DIMMs per channel population, the processor choice, and this BIOS option. Additional power savings can be achieved by reducing the memory frequency, at the expense of reduced performance. Read-only unless System Profile is set to Custom.

This field enables/disabled Efficiency Optimized Mode. Efficiency Optimized Mode maximizes Performance-per-Watt by opportunistically reducing frequency/power.

NUMA nodes per socket (NPS) field allows you to configure the memory NUMA domains per socket. The configuration can consist of one whole domain (NPS1), two domains (NPS2), or four domains (NPS4). In the case of a two-socket platform, an additional NPS profile is available to have whole system memory to be mapped as single NUMA domain (NPS0).

In addition to selecting the number of NUMA domains via NPS option, the processor allows for making memory per CCX as NUMA domain. In the processor each CCD has a maximum of two CCXs with each CCX having a shared last-level cache (LLC, or L3 cache) for all cores. The CCX as NUMA domain option allows for each LLC to be configured as a NUMA domain so that for certain workloads pinning execution to a single NUMA domain can be done.

When Adaptive Double DRAM Device Correction (ADDDC) is enabled, failing DRAM’s are dynamically mapped out. When set to enabled, it can have some impact to system performance under certain workloads. This feature is applicable for x4 DIMMs only.

Enables or disables Data Cache Unit (DCU) Streamer Prefetcher. This setting can affect performance, depending on the application running on the server. DCU streamer prefetchers detect multiple reads to a single cache line in a certain period of time and choose to load the following cache line to the L1 data caches. Recommended for High Performance Computing applications.

Enables or disables Data Cache Unit (DCU) IP Prefetcher. DCU IP Prefetcher looks for sequential load history to determine whether to prefetch the data to the L1 caches.

Governs the Boost Technology. This feature allows the processor cores to be automatically clocked up in frequency beyond the advertised processor speed. The amount of increased frequency (or 'turbo upside') one can expect from an EPYC processor depends on the fewer cores being exercised with work the higher the potential turbo upside. The potential drawback for Boost are mainly centered on increased power consumption and possible frequency jitter that can affect a small minority of latency-sensitive environments.

When set to Enabled, the processor is allowed to switch to minimum performance state when idle.

When Enabled, CPU interconnect bus link power management can reduce overall system power a bit while slightly reducing system performance.

Maximum Performance is typically selected for performance-centric workloads where it is acceptable to consume additional power to achieve the highest possible performance for the computing environment. This mode drives processor frequency to the maximum across all cores (although idled cores can still be frequency reduced by C-state enforcement through BIOS or OS mechanisms if enabled). This mode also offers the lowest latency of the CPU Power Management Mode options, so is always preferred.

The CPU uses the setting to manipulate the internal behavior of the processor and determines whether to target higher performance or better power savings. The possible settings are: Performance, Balanced Performance, Balanced Energy, Energy Efficient.

Permits Energy Efficient Turbo to be Enabled or Disabled.
Energy Efficient Turbo (EET) is a mode of operation where a processor's core frequency is adjusted within the turbo range based on workload.

Each processor core supports up to two logical processors. When set to Enabled, the BIOS reports all logical processors. When set to Disabled, the BIOS only reports one logical processor per core. Generally, higher processor count results in increased performance for most multi-threaded workloads and the recommendation is to keep this enabled. However, there are some floating point/scientific workloads, including HPC workloads, where disabling this feature may result in higher performance.

Patrol Scrubbing searches the memory for errors and repairs correctable errors to prevent the accumulation of memory errors. When set to Disabled, no patrol scrubbing will occur. When set to Standard Mode, the entire memory array will be scrubbed once in a 24 hour period. When set to Extended Mode, the entire memory array will be scrubbed more frequently to further increase system reliability.

The memory controller will periodically refresh the data in memory. The frequency at which memory is normally refreshed is referred to as 1X refresh rate. When memory modules are operating at a higher than normal temperature or to further increase system reliability, the refresh rate can be set to 2X, but may have a negative impact on memory subsystem performance under some circumstances.

When Enabled, PCIe Advanced State Power Management (ASPM) can reduce overall system power a bit while slightly reducing system performance.

NOTE: Some devices may not perform properly (they may hang or cause the system to hang) when ASPM is enabled, for this reason L1 will only be enabled for validated qualified cards.

When set to Custom, you can change setting of each option. Under Custom mode when C States is enabled, Monitor/Mwait should also be Enabled.

Specifies whether Monitor/Mwait instructions are enabled. Monitor/Mwait is only active when C States is set to Disabled.

When Enabled, Sub NUMA Clustering (SNC) is a feature for breaking up the LLC into disjoint clusters based on address range, with each cluster bound to a subset of the memory controllers in the system. It improves average latency to the LLC.

Selects the Processor Uncore Frequency.
Dynamic mode allows processor to optimize power resources across the cores and uncore during runtime. The optimization of the uncore frequency to either save power or optimize performance is influenced by the setting of the Energy Efficient Policy.

When set to Enabled, the BIOS will enable processor Virtualization features and provide the virtualization support to the Operating System (OS) through the DMAR table. In general, only virtualized environments such as VMware(r) ESX (tm), Microsoft Hyper-V(r) , Red Hat(r) KVM, and other virtualized operating systems will take advantage of these features. Disabling this feature is not known to significantly alter the performance or power characteristics of the system, so leaving this option Enabled is advised for most cases.

This kernel option sets adaptive tick mode (NOHZ_FULL) to specified processors. Since the number of interrupts is reduced to ones per second, latency-sensitive applications can take advantage of it.

When Enabled, memory interleaving is supported if a symmetric memory configuration is installed. When set to Disabled, the system supports Non-Uniform Memory Access (NUMA) (asymmetric) memory configurations. Channel interleaving is available with all configurations and is the intra-die memory interleave option. With channel interleaving the memory behind each UMC will be interleaved and seen as 1 NUMA domain per die.

For questions about the meanings of these flags, please contact the tester.
For other inquiries, please contact info@spec.org
Copyright 2017-2021 Standard Performance Evaluation Corporation
Tested with SPEC CPU2017 v1.1.5.
Report generated on 2021-03-30 15:32:51 by SPEC CPU2017 flags formatter v5178.

	Indicates that the flag description came from the user flags file.
	Indicates that the flag description came from the suite-wide flags file.
	Indicates that the flag description came from a per-benchmark flags file.

CPU2017 Flag DescriptionDell Inc. PowerEdge R6525 (AMD EPYC 7313 16-Core Processor)

Compilers: AMD Optimizing C/C++ Compiler Suite

Base Compiler Invocation

Peak Compiler Invocation

Base Portability Flags

Peak Portability Flags

Base Optimization Flags

Peak Optimization Flags

Base Other Flags

Peak Other Flags

Implicitly Included Flags

CPU2017 Flag Description
Dell Inc. PowerEdge R6525 (AMD EPYC 7313 16-Core Processor)