Example: biology

Intel® 64 and IA-32 Architectures Optimization Reference ...

intel 64 and IA-32 ArchitecturesOptimization Reference ManualOrder Number: 248966-046 AJanuary 2023 Notices & DisclaimersIntel technologies may require enabled hardware, software or service product or component can be absolutely secure. Your costs and results may vary. You may not use or facilitate the use of this document in connection with any infringement or other legal analysis concerning intel products described herein. You agree to grant intel a non-exclusive, royalty-free license to any patent claim thereafter drafted which includes subject matter disclosed product plans and roadmaps are subject to change without notice. The products described may contain design defects or errors known as errata which may cause the product to deviate from published specifications.

Intel® 64 and IA-32 Architectures Optimization Reference Manual Order Number: 248966-033 June 2016

Tags:

  Intel, Optimization

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Transcription of Intel® 64 and IA-32 Architectures Optimization Reference ...

1 intel 64 and IA-32 ArchitecturesOptimization Reference ManualOrder Number: 248966-046 AJanuary 2023 Notices & DisclaimersIntel technologies may require enabled hardware, software or service product or component can be absolutely secure. Your costs and results may vary. You may not use or facilitate the use of this document in connection with any infringement or other legal analysis concerning intel products described herein. You agree to grant intel a non-exclusive, royalty-free license to any patent claim thereafter drafted which includes subject matter disclosed product plans and roadmaps are subject to change without notice. The products described may contain design defects or errors known as errata which may cause the product to deviate from published specifications.

2 Current characterized errata are available on disclaims all express and implied warranties, including without limitation, the implied warranties of merchantability, fitness for a particular purpose, and non-infringement, as well as any warranty arising from course of performance, course of dealing, or usage in names are used by intel to identify products, technologies, or services that are in development and not publicly available. These are not commercial names and not intended to function as license (express or implied, by estoppel or otherwise) to any intellectual property rights is granted by this document, with the sole exception that a) you may publish an unmodified copy and b) code included in this document is licensed subject to the Zero-Clause BSD open source license (0 BSD), You may create software implementations based on this document and in compliance with the foregoing that are intended to execute on the intel product(s) referenced in this document.

3 No rights are granted to create modifications or derivatives of this document. intel Corporation. intel , the intel logo, and other intel marks are trademarks of intel Corporation or its subsid-iaries. Other names and brands may be claimed as the property of HISTORYDateRevisionDescriptionJanuary 2023046 Introduction to the 4th Generation intel Xeon Scalable Family ofProcessors Optimization of scalability and communications for the 4th GenerationIntel Xeon Scalable Family of Processors. intel Advanced Vector Extensions 512 - FP16 Instruction Set for intel Xeon Processors (Chapter 19) intel Advanced Matrix Extensions ( intel AMX) (Chapter 20) intel QuickAssist Technology (QAT) (Chapter 22) Update to Appendix BJanuary 2023046A Correction of author Correction of incorrect linkiiiCONTENTSPAGECHAPTER YOUR APPLICATION.

4 THIS MANUAL.. INFORMATION .. 1-4 CHAPTER 2 intel 64 AND IA-32 PROCESSOR RAPIDS MICROARCHITECTURE .. Generation intel Xeon Scalable Family of Processors .. LAKE PERFORMANCE HYBRID ARCHITECTURE.. Generation intel Core Processors Supporting Performance Hybrid Architecture .. Scheduling.. Thread Director .. with intel Hyper-Threading Technology Enabled on Processors Supporting x86 Hybrid Architecture.. with a Multi-E-Core Module .. Background Threads on x86 Hybrid Architecture .. for Application Developers .. COVE MICROARCHITECTURE .. Cove Microarchitecture Overview .. Front End .. Out-of-Order and Execution Engines.. Subsystem and Memory Subsystem.

5 Destination False Dependency.. LAKE CLIENT MICROARCHITECTURE .. Lake Client Microarchitecture Overview.. Front End .. Out of Order and Execution Engines .. and Memory Subsystem .. 2-17 Paired Stores .. Instructions .. Lake Client Microarchitecture Power Management .. SERVER MICROARCHITECTURE.. Server Microarchitecture Cache.. Mid-Level Cache .. Last Level Cache .. Server Microarchitecture Cache Recommendations .. Stores on Skylake Server Microarchitecture.. Server Power Management .. CLIENT MICROARCHITECTURE.. Front End .. Out-of-Order Execution Engine.. and Memory Subsystem .. Latency in Skylake Client Microarchitecture.. HYPER-THREADING TECHNOLOGY.

6 Resources and HT Technology .. Resources .. Resources .. Resources .. Pipeline and HT Technology .. Core .. TECHNOLOGY .. OF SIMD TECHNOLOGIES AND APPLICATION LEVEL EXTENSIONS .. Technology .. SIMD Extensions .. SIMD Extensions 2.. SIMD Extensions 3.. Streaming SIMD Extensions 3 .. and PCLMULQDQ .. Advanced Vector Extensions .. Floating-Point Conversion (F16C) .. (FMA) Extensions .. AVX2 .. Bit-Processing Instructions .. Transactional Synchronization Extensions.. and ADOX Instructions .. 2-43 CHAPTER 3 GENERAL Optimization TOOLS .. C++ and Fortran Compilers .. Compiler Recommendations .. Performance Analyzer .. PERSPECTIVES.

7 Dispatch Strategy and Compatible Code Strategy .. Cache-Parameter Strategy.. Strategy and Hardware Multithreading Support .. RULES, SUGGESTIONS AND TUNING HINTS .. THE FRONT END.. Prediction Optimization .. Branches.. Prediction .. , Calls and Returns .. Alignment.. Type Selection .. Unrolling .. and Decode Optimization .. for Micro-fusion.. for Macrofusion.. Prefixes (LCP) .. the Loop Stream Detector (LSD) .. for Decoded ICache .. Decoding Guidelines .. THE EXECUTION CORE .. Selection .. Divide .. LEA .. and SBB in Sandy Bridge Microarchitecture .. Rotation.. Bit Count Rotation and Shift .. Calculations .. Registers and Dependency Breaking Idioms.

8 NOPs.. SIMD Data Types.. Scheduling .. MOV Instructions.. Stalls in Execution Core .. Bus Conflicts .. between Execution Domains.. Register Stalls .. XMM Register Stalls.. Flag Register Stalls.. Operands.. of Partially Vectorizable Code.. Packing Techniques .. Result Passing.. Optimization .. Considerations .. MEMORY ACCESSES.. and Store Execution Bandwidth .. Use of Load Bandwidth in Sandy Bridge Microarchitecture.. Cache Latency in Sandy Bridge Microarchitecture .. L1D Cache Bank Conflict .. Register Spills .. Speculative Execution and Memory Disambiguation.. Forwarding .. Restriction on Size and Alignment .. Restriction on Data Availability.

9 Layout Optimizations .. Alignment .. Limits and Aliasing in Caches.. Code and Data .. Code .. Independent Code .. Combining .. Enhancement.. Store Bus Traffic .. Instruction Fetching and Software Prefetching.. Prefetching for First-Level Data Cache.. Prefetching for Second-Level Cache .. Instructions .. Prefix and Data Movement .. REP MOVSB and STOSB Operation .. Short REP MOVSB.. Considerations .. Considerations .. Considerations .. STRING OPERATIONS.. Zero Length REP MOVSB .. Short REP STOSB .. Short REP CMPSB and SCASB .. CONSIDERATIONS .. for Optimizing Floating-point Code .. Modes and Exceptions .. Exceptions .. with floating-point exceptions in x87 FPU code.

10 Exceptions in SSE/SSE2/SSE3 Code.. Modes.. Mode .. vs. Scalar SIMD Floating-point Trade-offs .. SSE/SSE2 .. Functions.. MAXIMIZING PCIE PERFORMANCE.. PCIe Performance for Accesses Toward Coherent Memory and Toward MMIO Regions (P2P) .. SCALABILITY WITH CONTENDED LINE ACCESS IN 4TH GENERATION intel XEON SCALABLE PROCESSORS .. it Happens.. to Detect it.. to Fix it .. Study: SysBench/MariaDB Metric CHA % Cycles Fast Asserted .. Sequence Slowdowns.. it Happens.. to Detect it.. to Fix it .. for Branches >2GB .. it Happens.. to Detect it.. to Fix it .. OPTIMIZING COMMUNICATION WITH PCI DEVICES ON 4TH GENERATION intel XEON SCALABLE PROCESSORS.


Related search queries