diff --git a/README.md b/README.md index 851af9980a5..7763d17f2f4 100644 --- a/README.md +++ b/README.md @@ -74,7 +74,7 @@ The OpenVINO™ Runtime can infer models on different hardware devices. This sec ARM CPU openvino_arm_cpu_plugin - Raspberry Pi™ 4 Model B, Apple® Mac mini with Apple silicon + ARM CPUs with armv7a and higher, ARM64 CPUs with arm64-v8a and higher, Apple® Mac with Apple silicon GPU diff --git a/docs/articles_en/about-openvino/release-notes-openvino/system-requirements.rst b/docs/articles_en/about-openvino/release-notes-openvino/system-requirements.rst index bb0e45b5a96..b1bffb3251c 100644 --- a/docs/articles_en/about-openvino/release-notes-openvino/system-requirements.rst +++ b/docs/articles_en/about-openvino/release-notes-openvino/system-requirements.rst @@ -26,7 +26,7 @@ CPU * 6th - 14th generation Intel® Core™ processors * Intel® Core™ Ultra (codename Meteor Lake) * 1st - 5th generation Intel® Xeon® Scalable Processors - * ARM and ARM64 CPUs; Apple M1, M2, and Raspberry Pi + * ARM CPUs with armv7a and higher, ARM64 CPUs with arm64-v8a and higher, Apple® Mac with Apple silicon .. tab-item:: Supported Operating Systems diff --git a/docs/articles_en/openvino-workflow/running-inference/inference-devices-and-modes/cpu-device.rst b/docs/articles_en/openvino-workflow/running-inference/inference-devices-and-modes/cpu-device.rst index 46f33ee3922..ede4765b67d 100644 --- a/docs/articles_en/openvino-workflow/running-inference/inference-devices-and-modes/cpu-device.rst +++ b/docs/articles_en/openvino-workflow/running-inference/inference-devices-and-modes/cpu-device.rst @@ -14,7 +14,7 @@ CPU Device The CPU plugin is a part of the Intel® Distribution of OpenVINO™ toolkit. It is developed to achieve high performance inference of neural networks on Intel® x86-64 and Arm® CPUs. The newer 11th generation and later Intel® CPUs provide even further performance boost, especially with INT8 models. For an in-depth description of CPU plugin, see: -- `CPU plugin developer documentation `__. +- `CPU plugin developer documentation `__. - `OpenVINO Runtime CPU plugin source files `__. .. note:: @@ -60,6 +60,7 @@ CPU plugin supports the following data types as inference precision of internal - ``f32`` (Intel® x86-64, Arm®) - ``bf16`` (Intel® x86-64) + - ``f16`` (Intel® x86-64, Arm®) - Integer data types: - ``i32`` (Intel® x86-64, Arm®) @@ -92,21 +93,24 @@ CPU plugin supports the following floating-point data types as inference precisi - ``f32`` (Intel® x86-64, Arm®) - ``bf16`` (Intel® x86-64) +- ``f16`` (Intel® x86-64, Arm®) -The default floating-point precision of a CPU primitive is ``f32``. To support the ``f16`` OpenVINO IR the plugin internally converts -all the ``f16`` values to ``f32`` and all the calculations are performed using the native precision of ``f32``. -On platforms that natively support ``bfloat16`` calculations (have the ``AVX512_BF16`` or ``AMX`` extension), the ``bf16`` type is automatically used instead +The default floating-point precision of a CPU primitive is ``f32``. To support the ``f16`` OpenVINO IR on platforms that do not natively support ``float16``, the plugin internally converts +all the ``f16`` values to ``f32``, and all calculations are performed using the native precision of ``f32``. +On platforms that natively support half-precision calculations (``bfloat16`` or ``float16``), the half-precision type (``bf16`` or ``f16``) is automatically used instead of ``f32`` to achieve better performance (see the `Execution Mode Hint <#execution-mode-hint>`__). -Thus, no special steps are required to run a ``bf16`` model. For more details about the ``bfloat16`` format, see +Thus, no special steps are required to run a model with ``bf16`` or ``f16`` inference precision. + +Using the half-precision provides the following performance benefits: + +- ``bfloat16`` and ``float16`` data types enable Intel® Advanced Matrix Extension (AMX) on 4+ generation Intel® Xeon® Scalable Processors, resulting in significantly faster computations on the corresponding hardware compared to AVX512 or AVX2 instructions in many deep learning operation implementations. +- ``float16`` data type enables the ``armv8.2-a+fp16`` extension on ARM64 CPUs, which significantly improves performance due to the doubled vector capacity. +- Memory footprint is reduced since most weight and activation tensors are stored in half-precision. + +For more details about the ``bfloat16`` format, see the `BFLOAT16 – Hardware Numerics Definition white paper `__. - -Using the ``bf16`` precision provides the following performance benefits: - -- ``bfloat16`` data type allows using Intel® Advanced Matrix Extension (AMX), which provides dramatically faster computations on corresponding hardware in comparison with AVX512 or AVX2 instructions in many DL operation implementations. -- Reduced memory consumption since ``bfloat16`` data half the size of 32-bit float. - -To check if the CPU device can support the ``bfloat16`` data type, use the :doc:`query device properties interface ` -to query ``ov::device::capabilities`` property, which should contain ``BF16`` in the list of CPU capabilities: +To check if the CPU device can support the half-precision data type, use the :doc:`query device properties interface ` +to query ``ov::device::capabilities`` property, which should contain ``FP16`` or ``BF16`` in the list of CPU capabilities: .. tab-set:: @@ -129,7 +133,7 @@ to query ``ov::device::capabilities`` property, which should contain ``BF16`` in Inference Precision Hint ----------------------------------------------------------- -If the model has been converted to ``bf16``, the ``ov::hint::inference_precision`` is set to ``ov::element::bf16`` and can be checked via +If the model has been converted to half-precision (``bf16`` or ``f16``), the ``ov::hint::inference_precision`` is set to ``ov::element::f16`` or ``ov::element::bf16`` and can be checked via the ``ov::CompiledModel::get_property`` call. The code below demonstrates how to get the element type: .. tab-set:: @@ -148,7 +152,7 @@ the ``ov::CompiledModel::get_property`` call. The code below demonstrates how to :language: cpp :fragment: [part1] -To infer the model in ``f32`` precision instead of ``bf16`` on targets with native ``bf16`` support, set the ``ov::hint::inference_precision`` to ``ov::element::f32``. +To infer the model in ``f32`` precision instead of half-precision (``bf16`` or ``f16``) on targets with native half-precision support, set the ``ov::hint::inference_precision`` to ``ov::element::f32``. .. tab-set:: @@ -178,16 +182,16 @@ To enable the simulation, the ``ov::hint::inference_precision`` has to be explic .. note:: - Due to the reduced mantissa size of the ``bfloat16`` data type, the resulting ``bf16`` inference accuracy may differ from the ``f32`` inference, - especially for models that were not trained using the ``bfloat16`` data type. If the ``bf16`` inference accuracy is not acceptable, + Due to the reduced mantissa size of half-precision data types (``bfloat16`` or ``float16``), the resulting half-precision inference accuracy may differ from the ``f32`` inference, + especially for models that were not trained using half-precision data types. If half-precision inference accuracy is not acceptable, it is recommended to switch to the ``f32`` precision. Also, the performance/accuracy balance can be managed using the ``ov::hint::execution_mode`` hint, see the `Execution Mode Hint <#execution-mode-hint>`__. Execution Mode Hint ----------------------------------------------------------- In case ``ov::hint::inference_precision`` is not explicitly set, one can use ``ov::hint::execution_mode`` hint to direct the run-time optimizations toward either better accuracy or better performance. -If ``ov::hint::execution_mode`` is set to ``ov::hint::ExecutionMode::PERFORMANCE`` (default behavior) and the platform natively supports ``bfloat16`` -calculations (has the ``AVX512_BF16`` or ``AMX`` extension) then ``bf16`` type is automatically used instead of ``f32`` to achieve better performance. +If ``ov::hint::execution_mode`` is set to ``ov::hint::ExecutionMode::PERFORMANCE`` (default behavior) and the platform natively supports half-precision +calculations (``bfloat16`` or ``float16``) then ``bf16`` or ``f16`` type is automatically used instead of ``f32`` to achieve better performance. If the accuracy in this mode is not good enough, then set ``ov::hint::execution_mode`` to ``ov::hint::ExecutionMode::ACCURACY`` to enforce the plugin to use the ``f32`` precision in floating point calculations. @@ -237,10 +241,6 @@ For more details, see the :doc:`optimization guide <../optimize-inference>`. on data transfer between NUMA nodes. In that case it is better to use the ``ov::hint::PerformanceMode::LATENCY`` performance hint. For more details see the :doc:`performance hints <../optimize-inference/high-level-performance-hints>` overview. -.. note:: - - Multi-stream execution is not supported on Arm® platforms. Latency and throughput hints have identical behavior and use only one stream for inference. - Dynamic Shapes +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ diff --git a/docs/dev/build_raspbian.md b/docs/dev/build_raspbian.md index 3fcf07f2cdd..0a7dac5a43b 100644 --- a/docs/dev/build_raspbian.md +++ b/docs/dev/build_raspbian.md @@ -3,7 +3,7 @@ > **NOTE**: Since 2023.0 release, you can compile [OpenVINO Intel CPU plugin](https://github.com/openvinotoolkit/openvino/tree/master/src/plugins/intel_cpu) on ARM platforms. ## Hardware Requirements -* Raspberry Pi 2 or 3 with Raspbian Stretch OS (32 or 64-bit). +* Raspberry Pi with Raspbian Stretch OS or Raspberry Pi OS (32 or 64-bit). > **NOTE**: Despite the Raspberry Pi CPU is ARMv8, 32-bit OS detects ARMv7 CPU instruction set. The default `gcc` compiler applies ARMv6 architecture flag for compatibility with lower versions of boards. For more information, run the `gcc -Q --help=target` command and refer to the description of the `-march=` option.