diff --git a/README.md b/README.md
index 851af9980a5..7763d17f2f4 100644
--- a/README.md
+++ b/README.md
@@ -74,7 +74,7 @@ The OpenVINO™ Runtime can infer models on different hardware devices. This sec
| ARM CPU
| openvino_arm_cpu_plugin |
- Raspberry Pi™ 4 Model B, Apple® Mac mini with Apple silicon
+ | ARM CPUs with armv7a and higher, ARM64 CPUs with arm64-v8a and higher, Apple® Mac with Apple silicon
|
| GPU |
diff --git a/docs/articles_en/about-openvino/release-notes-openvino/system-requirements.rst b/docs/articles_en/about-openvino/release-notes-openvino/system-requirements.rst
index bb0e45b5a96..b1bffb3251c 100644
--- a/docs/articles_en/about-openvino/release-notes-openvino/system-requirements.rst
+++ b/docs/articles_en/about-openvino/release-notes-openvino/system-requirements.rst
@@ -26,7 +26,7 @@ CPU
* 6th - 14th generation Intel® Core™ processors
* Intel® Core™ Ultra (codename Meteor Lake)
* 1st - 5th generation Intel® Xeon® Scalable Processors
- * ARM and ARM64 CPUs; Apple M1, M2, and Raspberry Pi
+ * ARM CPUs with armv7a and higher, ARM64 CPUs with arm64-v8a and higher, Apple® Mac with Apple silicon
.. tab-item:: Supported Operating Systems
diff --git a/docs/articles_en/openvino-workflow/running-inference/inference-devices-and-modes/cpu-device.rst b/docs/articles_en/openvino-workflow/running-inference/inference-devices-and-modes/cpu-device.rst
index 46f33ee3922..ede4765b67d 100644
--- a/docs/articles_en/openvino-workflow/running-inference/inference-devices-and-modes/cpu-device.rst
+++ b/docs/articles_en/openvino-workflow/running-inference/inference-devices-and-modes/cpu-device.rst
@@ -14,7 +14,7 @@ CPU Device
The CPU plugin is a part of the Intel® Distribution of OpenVINO™ toolkit. It is developed to achieve high performance inference of neural networks on Intel® x86-64 and Arm® CPUs. The newer 11th generation and later Intel® CPUs provide even further performance boost, especially with INT8 models.
For an in-depth description of CPU plugin, see:
-- `CPU plugin developer documentation `__.
+- `CPU plugin developer documentation `__.
- `OpenVINO Runtime CPU plugin source files `__.
.. note::
@@ -60,6 +60,7 @@ CPU plugin supports the following data types as inference precision of internal
- ``f32`` (Intel® x86-64, Arm®)
- ``bf16`` (Intel® x86-64)
+ - ``f16`` (Intel® x86-64, Arm®)
- Integer data types:
- ``i32`` (Intel® x86-64, Arm®)
@@ -92,21 +93,24 @@ CPU plugin supports the following floating-point data types as inference precisi
- ``f32`` (Intel® x86-64, Arm®)
- ``bf16`` (Intel® x86-64)
+- ``f16`` (Intel® x86-64, Arm®)
-The default floating-point precision of a CPU primitive is ``f32``. To support the ``f16`` OpenVINO IR the plugin internally converts
-all the ``f16`` values to ``f32`` and all the calculations are performed using the native precision of ``f32``.
-On platforms that natively support ``bfloat16`` calculations (have the ``AVX512_BF16`` or ``AMX`` extension), the ``bf16`` type is automatically used instead
+The default floating-point precision of a CPU primitive is ``f32``. To support the ``f16`` OpenVINO IR on platforms that do not natively support ``float16``, the plugin internally converts
+all the ``f16`` values to ``f32``, and all calculations are performed using the native precision of ``f32``.
+On platforms that natively support half-precision calculations (``bfloat16`` or ``float16``), the half-precision type (``bf16`` or ``f16``) is automatically used instead
of ``f32`` to achieve better performance (see the `Execution Mode Hint <#execution-mode-hint>`__).
-Thus, no special steps are required to run a ``bf16`` model. For more details about the ``bfloat16`` format, see
+Thus, no special steps are required to run a model with ``bf16`` or ``f16`` inference precision.
+
+Using the half-precision provides the following performance benefits:
+
+- ``bfloat16`` and ``float16`` data types enable Intel® Advanced Matrix Extension (AMX) on 4+ generation Intel® Xeon® Scalable Processors, resulting in significantly faster computations on the corresponding hardware compared to AVX512 or AVX2 instructions in many deep learning operation implementations.
+- ``float16`` data type enables the ``armv8.2-a+fp16`` extension on ARM64 CPUs, which significantly improves performance due to the doubled vector capacity.
+- Memory footprint is reduced since most weight and activation tensors are stored in half-precision.
+
+For more details about the ``bfloat16`` format, see
the `BFLOAT16 – Hardware Numerics Definition white paper `__.
-
-Using the ``bf16`` precision provides the following performance benefits:
-
-- ``bfloat16`` data type allows using Intel® Advanced Matrix Extension (AMX), which provides dramatically faster computations on corresponding hardware in comparison with AVX512 or AVX2 instructions in many DL operation implementations.
-- Reduced memory consumption since ``bfloat16`` data half the size of 32-bit float.
-
-To check if the CPU device can support the ``bfloat16`` data type, use the :doc:`query device properties interface `
-to query ``ov::device::capabilities`` property, which should contain ``BF16`` in the list of CPU capabilities:
+To check if the CPU device can support the half-precision data type, use the :doc:`query device properties interface `
+to query ``ov::device::capabilities`` property, which should contain ``FP16`` or ``BF16`` in the list of CPU capabilities:
.. tab-set::
@@ -129,7 +133,7 @@ to query ``ov::device::capabilities`` property, which should contain ``BF16`` in
Inference Precision Hint
-----------------------------------------------------------
-If the model has been converted to ``bf16``, the ``ov::hint::inference_precision`` is set to ``ov::element::bf16`` and can be checked via
+If the model has been converted to half-precision (``bf16`` or ``f16``), the ``ov::hint::inference_precision`` is set to ``ov::element::f16`` or ``ov::element::bf16`` and can be checked via
the ``ov::CompiledModel::get_property`` call. The code below demonstrates how to get the element type:
.. tab-set::
@@ -148,7 +152,7 @@ the ``ov::CompiledModel::get_property`` call. The code below demonstrates how to
:language: cpp
:fragment: [part1]
-To infer the model in ``f32`` precision instead of ``bf16`` on targets with native ``bf16`` support, set the ``ov::hint::inference_precision`` to ``ov::element::f32``.
+To infer the model in ``f32`` precision instead of half-precision (``bf16`` or ``f16``) on targets with native half-precision support, set the ``ov::hint::inference_precision`` to ``ov::element::f32``.
.. tab-set::
@@ -178,16 +182,16 @@ To enable the simulation, the ``ov::hint::inference_precision`` has to be explic
.. note::
- Due to the reduced mantissa size of the ``bfloat16`` data type, the resulting ``bf16`` inference accuracy may differ from the ``f32`` inference,
- especially for models that were not trained using the ``bfloat16`` data type. If the ``bf16`` inference accuracy is not acceptable,
+ Due to the reduced mantissa size of half-precision data types (``bfloat16`` or ``float16``), the resulting half-precision inference accuracy may differ from the ``f32`` inference,
+ especially for models that were not trained using half-precision data types. If half-precision inference accuracy is not acceptable,
it is recommended to switch to the ``f32`` precision. Also, the performance/accuracy balance can be managed using the ``ov::hint::execution_mode`` hint,
see the `Execution Mode Hint <#execution-mode-hint>`__.
Execution Mode Hint
-----------------------------------------------------------
In case ``ov::hint::inference_precision`` is not explicitly set, one can use ``ov::hint::execution_mode`` hint to direct the run-time optimizations toward either better accuracy or better performance.
-If ``ov::hint::execution_mode`` is set to ``ov::hint::ExecutionMode::PERFORMANCE`` (default behavior) and the platform natively supports ``bfloat16``
-calculations (has the ``AVX512_BF16`` or ``AMX`` extension) then ``bf16`` type is automatically used instead of ``f32`` to achieve better performance.
+If ``ov::hint::execution_mode`` is set to ``ov::hint::ExecutionMode::PERFORMANCE`` (default behavior) and the platform natively supports half-precision
+calculations (``bfloat16`` or ``float16``) then ``bf16`` or ``f16`` type is automatically used instead of ``f32`` to achieve better performance.
If the accuracy in this mode is not good enough, then set ``ov::hint::execution_mode`` to ``ov::hint::ExecutionMode::ACCURACY`` to enforce the plugin to
use the ``f32`` precision in floating point calculations.
@@ -237,10 +241,6 @@ For more details, see the :doc:`optimization guide <../optimize-inference>`.
on data transfer between NUMA nodes. In that case it is better to use the ``ov::hint::PerformanceMode::LATENCY`` performance hint.
For more details see the :doc:`performance hints <../optimize-inference/high-level-performance-hints>` overview.
-.. note::
-
- Multi-stream execution is not supported on Arm® platforms. Latency and throughput hints have identical behavior and use only one stream for inference.
-
Dynamic Shapes
+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
diff --git a/docs/dev/build_raspbian.md b/docs/dev/build_raspbian.md
index 3fcf07f2cdd..0a7dac5a43b 100644
--- a/docs/dev/build_raspbian.md
+++ b/docs/dev/build_raspbian.md
@@ -3,7 +3,7 @@
> **NOTE**: Since 2023.0 release, you can compile [OpenVINO Intel CPU plugin](https://github.com/openvinotoolkit/openvino/tree/master/src/plugins/intel_cpu) on ARM platforms.
## Hardware Requirements
-* Raspberry Pi 2 or 3 with Raspbian Stretch OS (32 or 64-bit).
+* Raspberry Pi with Raspbian Stretch OS or Raspberry Pi OS (32 or 64-bit).
> **NOTE**: Despite the Raspberry Pi CPU is ARMv8, 32-bit OS detects ARMv7 CPU instruction set. The default `gcc` compiler applies ARMv6 architecture flag for compatibility with lower versions of boards. For more information, run the `gcc -Q --help=target` command and refer to the description of the `-march=` option.