From 0aa5a8f704683ba95f2cbc797237656747c02bc7 Mon Sep 17 00:00:00 2001 From: Sebastian Golebiewski Date: Mon, 21 Aug 2023 16:47:28 +0200 Subject: [PATCH] port-19307 (#19310) Porting: https://github.com/openvinotoolkit/openvino/pull/19307 Updating tutorials: adding table of contents and new notebooks. --- docs/nbdoc/consts.py | 2 +- .../notebooks/001-hello-world-with-output.rst | 39 +- .../index.html | 6 +- .../002-openvino-api-with-output.rst | 112 +- .../003-hello-segmentation-with-output.rst | 60 +- .../index.html | 10 +- .../004-hello-detection-with-output.rst | 44 +- .../index.html | 8 +- ...classification-to-openvino-with-output.rst | 135 +- .../index.html | 6 +- ...2-pytorch-onnx-to-openvino-with-output.rst | 133 +- .../index.html | 10 +- .../102-pytorch-to-openvino-with-output.rst | 182 +- .../index.html | 32 +- ...to-openvino-classification-with-output.rst | 93 +- .../index.html | 14 +- .../notebooks/104-model-tools-with-output.rst | 126 +- ...105-language-quantize-bert-with-output.rst | 129 +- .../notebooks/106-auto-device-with-output.rst | 162 +- .../106-auto-device-with-output_25_0.png | 4 +- .../106-auto-device-with-output_26_0.png | 4 +- .../index.html | 12 +- ...tion-quantization-data2vec-with-output.rst | 78 +- docs/notebooks/108-gpu-device-with-output.rst | 1375 +++++++++++++++ .../109-latency-tricks-with-output.rst | 145 +- .../109-latency-tricks-with-output_30_0.png | 4 +- .../index.html | 20 +- .../109-throughput-tricks-with-output.rst | 717 ++++++++ ...109-throughput-tricks-with-output_14_0.jpg | 3 + ...109-throughput-tricks-with-output_17_0.jpg | 3 + ...109-throughput-tricks-with-output_20_0.jpg | 3 + ...109-throughput-tricks-with-output_22_0.jpg | 3 + ...109-throughput-tricks-with-output_26_0.jpg | 3 + ...109-throughput-tricks-with-output_28_0.jpg | 3 + ...109-throughput-tricks-with-output_30_0.jpg | 3 + ...109-throughput-tricks-with-output_33_0.png | 3 + .../109-throughput-tricks-with-output_4_0.jpg | 3 + .../index.html | 15 + ...110-ct-scan-live-inference-with-output.rst | 126 +- .../index.html | 6 +- ...segmentation-quantize-nncf-with-output.rst | 260 +-- ...ntation-quantize-nncf-with-output_37_1.png | 4 +- .../index.html | 10 +- ...ov5-quantization-migration-with-output.rst | 336 ++-- ...antization-migration-with-output_34_0.png} | 0 ...antization-migration-with-output_40_0.png} | 0 .../index.html | 8 +- ...training-quantization-nncf-with-output.rst | 226 +-- ...lassification-quantization-with-output.rst | 164 +- ...ication-quantization-with-output_29_2.png} | 0 .../index.html | 6 +- docs/notebooks/115-async-api-with-output.rst | 155 +- .../115-async-api-with-output_21_0.png | 3 + .../index.html | 12 +- .../116-sparsity-optimization-with-output.rst | 109 +- .../117-model-server-with-output.rst | 159 +- .../index.html | 8 +- ...118-optimize-preprocessing-with-output.rst | 256 ++- ...timize-preprocessing-with-output_13_1.png} | 0 .../index.html | 6 +- .../119-tflite-to-openvino-with-output.rst | 122 +- .../index.html | 8 +- ...ject-detection-to-openvino-with-output.rst | 150 +- ...detection-to-openvino-with-output_38_0.png | 4 +- .../index.html | 8 +- .../121-convert-to-openvino-with-output.rst | 1474 +++++++++++++++++ .../201-vision-monodepth-with-output.rst | 116 +- .../index.html | 6 +- ...sion-superresolution-image-with-output.rst | 143 +- .../index.html | 12 +- ...sion-superresolution-video-with-output.rst | 209 ++- .../203-meter-reader-with-output.rst | 88 +- .../index.html | 14 +- ...nter-semantic-segmentation-with-output.rst | 187 ++- .../index.html | 12 +- ...-vision-background-removal-with-output.rst | 161 +- ...on-background-removal-with-output_20_0.png | 3 - ...on-background-removal-with-output_22_0.png | 4 +- ...on-background-removal-with-output_24_0.png | 3 + .../index.html | 8 +- ...206-vision-paddlegan-anime-with-output.rst | 193 ++- ...sion-paddlegan-anime-with-output_37_0.png} | 0 .../index.html | 8 +- ...-paddlegan-superresolution-with-output.rst | 132 +- ...egan-superresolution-with-output_27_1.png} | 0 ...egan-superresolution-with-output_31_1.png} | 0 ...egan-superresolution-with-output_33_0.png} | 0 .../index.html | 10 +- ...ical-character-recognition-with-output.rst | 177 +- ...haracter-recognition-with-output_15_0.png} | 0 ...character-recognition-with-output_23_0.png | 3 - ...character-recognition-with-output_25_0.png | 4 +- ...haracter-recognition-with-output_27_0.jpg} | 0 ...character-recognition-with-output_27_0.png | 3 + ...aracter-recognition-with-output_27_10.jpg} | 0 ...aracter-recognition-with-output_27_10.png} | 0 ...haracter-recognition-with-output_27_2.jpg} | 0 ...haracter-recognition-with-output_27_2.png} | 0 ...haracter-recognition-with-output_27_4.jpg} | 0 ...haracter-recognition-with-output_27_4.png} | 0 ...haracter-recognition-with-output_27_6.jpg} | 0 ...haracter-recognition-with-output_27_6.png} | 0 ...haracter-recognition-with-output_27_8.jpg} | 0 ...haracter-recognition-with-output_27_8.png} | 0 .../index.html | 32 +- .../209-handwritten-ocr-with-output.rst | 130 +- ... 209-handwritten-ocr-with-output_21_0.png} | 0 ... 209-handwritten-ocr-with-output_30_1.png} | 0 .../index.html | 8 +- ...slowfast-video-recognition-with-output.rst | 90 +- .../211-speech-to-text-with-output.rst | 176 +- .../index.html | 12 +- ...annote-speaker-diarization-with-output.rst | 119 +- ...-speaker-diarization-with-output_27_0.png} | 0 .../index.html | 10 +- .../213-question-answering-with-output.rst | 137 +- .../214-grammar-correction-with-output.rst | 149 +- .../215-image-inpainting-with-output.rst | 118 +- .../215-image-inpainting-with-output_12_0.png | 3 - .../215-image-inpainting-with-output_14_0.png | 4 +- .../215-image-inpainting-with-output_16_0.png | 4 +- .../215-image-inpainting-with-output_18_0.png | 3 + .../215-image-inpainting-with-output_20_0.png | 3 - .../215-image-inpainting-with-output_22_0.png | 3 + .../index.html | 12 +- .../216-attention-center-with-output.rst | 113 +- ...216-attention-center-with-output_14_1.png} | 0 ...216-attention-center-with-output_16_1.png} | 0 .../index.html | 8 +- .../217-vision-deblur-with-output.rst | 146 +- ...=> 217-vision-deblur-with-output_24_0.png} | 0 .../217-vision-deblur-with-output_25_0.png | 3 - .../217-vision-deblur-with-output_27_0.png | 4 +- .../217-vision-deblur-with-output_29_0.png | 3 + .../index.html | 10 +- ...-detection-and-recognition-with-output.rst | 125 +- ...tion-and-recognition-with-output_13_0.png} | 0 ...tion-and-recognition-with-output_20_0.png} | 0 ...tion-and-recognition-with-output_26_0.png} | 0 .../index.html | 10 +- ...219-knowledge-graphs-conve-with-output.rst | 272 +-- ...ss-lingual-books-alignment-with-output.rst | 1075 ++++++++++++ ...ngual-books-alignment-with-output_31_0.png | 3 + ...ngual-books-alignment-with-output_48_0.png | 3 + .../index.html | 8 + .../221-machine-translation-with-output.rst | 119 +- ...-vision-image-colorization-with-output.rst | 118 +- ...n-image-colorization-with-output_20_0.png} | 0 ...n-image-colorization-with-output_21_0.png} | 0 .../index.html | 8 +- .../223-text-prediction-with-output.rst | 253 +-- ...-segmentation-point-clouds-with-output.rst | 108 +- ...ntation-point-clouds-with-output_15_0.png} | 0 .../index.html | 8 +- ...le-diffusion-text-to-image-with-output.rst | 149 +- ...ffusion-text-to-image-with-output_29_1.png | 3 - ...ffusion-text-to-image-with-output_33_1.png | 4 +- ...ffusion-text-to-image-with-output_37_1.png | 3 + ...fusion-text-to-image-with-output_39_1.png} | 0 .../index.html | 10 +- .../226-yolov7-optimization-with-output.rst | 378 +++-- ...6-yolov7-optimization-with-output_22_0.jpg | 3 - ...6-yolov7-optimization-with-output_22_0.png | 3 - ...6-yolov7-optimization-with-output_26_0.jpg | 3 + ...6-yolov7-optimization-with-output_26_0.png | 3 + ...6-yolov7-optimization-with-output_38_0.jpg | 3 - ...6-yolov7-optimization-with-output_38_0.png | 3 - ...6-yolov7-optimization-with-output_43_0.jpg | 3 + ...6-yolov7-optimization-with-output_43_0.png | 3 + ...26-yolov7-optimization-with-output_8_0.jpg | 3 - ...26-yolov7-optimization-with-output_8_0.png | 3 - ...26-yolov7-optimization-with-output_9_0.jpg | 3 + ...26-yolov7-optimization-with-output_9_0.png | 3 + .../index.html | 16 +- ...isper-subtitles-generation-with-output.rst | 198 ++- ...28-clip-zero-shot-convert-with-output.rst} | 201 ++- ...lip-zero-shot-convert-with-output_12_0.png | 3 + ...lip-zero-shot-convert-with-output_17_0.png | 3 + ...clip-zero-shot-convert-with-output_4_0.png | 3 + .../index.html | 9 + ...-image-classification-with-output_10_0.png | 3 - ...-image-classification-with-output_15_0.png | 3 - ...t-image-classification-with-output_5_0.png | 3 - .../index.html | 9 - ...28-clip-zero-shot-quantize-with-output.rst | 376 +++++ ...ip-zero-shot-quantize-with-output_16_0.png | 3 + .../index.html | 7 + ...rt-sequence-classification-with-output.rst | 105 +- .../230-yolov8-optimization-with-output.rst | 761 +++++---- ...-yolov8-optimization-with-output_13_1.png} | 0 ...0-yolov8-optimization-with-output_14_1.png | 3 - ...0-yolov8-optimization-with-output_15_1.png | 3 + ...-yolov8-optimization-with-output_27_0.png} | 0 ...-yolov8-optimization-with-output_29_0.png} | 0 ...0-yolov8-optimization-with-output_54_0.png | 3 - ...0-yolov8-optimization-with-output_56_0.png | 3 - ...0-yolov8-optimization-with-output_59_0.png | 3 + ...0-yolov8-optimization-with-output_61_0.png | 3 + ...0-yolov8-optimization-with-output_85_0.png | 3 - ...0-yolov8-optimization-with-output_91_0.png | 3 + ...0-yolov8-optimization-with-output_97_0.png | 3 + ...0-yolov8-optimization-with-output_99_0.png | 3 + .../index.html | 20 +- ...ruct-pix2pix-image-editing-with-output.rst | 97 +- ...ix2pix-image-editing-with-output_25_0.png} | 0 .../index.html | 6 +- ...clip-language-saliency-map-with-output.rst | 115 +- ...language-saliency-map-with-output_15_0.png | 4 +- ...language-saliency-map-with-output_17_0.png | 4 +- ...language-saliency-map-with-output_19_1.png | 4 +- ...language-saliency-map-with-output_28_1.png | 3 - ...language-saliency-map-with-output_31_1.png | 3 + ...language-saliency-map-with-output_34_1.png | 3 - ...language-saliency-map-with-output_37_1.png | 3 + .../index.html | 11 - ...visual-language-processing-with-output.rst | 378 +++-- ...-language-processing-with-output_28_0.png} | 0 ...-language-processing-with-output_30_0.png} | 0 ...l-language-processing-with-output_8_0.png} | 0 .../index.html | 10 +- ...-encodec-audio-compression-with-output.rst | 127 +- ...ec-audio-compression-with-output_38_1.png} | 0 .../index.html | 10 +- ...ontrolnet-stable-diffusion-with-output.rst | 258 ++- ...lnet-stable-diffusion-with-output_14_0.png | 3 - ...lnet-stable-diffusion-with-output_17_0.png | 3 + ...olnet-stable-diffusion-with-output_8_0.png | 4 +- .../index.html | 8 +- ...diffusion-v2-infinite-zoom-with-output.rst | 323 +++- ...v2-optimum-demo-comparison-with-output.rst | 33 +- .../index.html | 10 +- ...-diffusion-v2-optimum-demo-with-output.rst | 50 +- .../index.html | 6 +- ...sion-v2-text-to-image-demo-with-output.rst | 101 +- ...2-text-to-image-demo-with-output_25_0.png} | 0 .../index.html | 6 +- ...diffusion-v2-text-to-image-with-output.rst | 315 ++-- .../237-segment-anything-with-output.rst | 855 ++++++++-- .../237-segment-anything-with-output_17_0.png | 3 - .../237-segment-anything-with-output_21_0.png | 3 + .../237-segment-anything-with-output_24_0.png | 3 - .../237-segment-anything-with-output_28_0.png | 3 + .../237-segment-anything-with-output_31_0.png | 3 - .../237-segment-anything-with-output_35_0.png | 4 +- .../237-segment-anything-with-output_39_0.png | 3 + .../237-segment-anything-with-output_40_0.png | 3 - .../237-segment-anything-with-output_44_0.png | 4 +- .../237-segment-anything-with-output_48_0.png | 3 + .../237-segment-anything-with-output_49_0.png | 3 - .../237-segment-anything-with-output_53_0.png | 3 + .../237-segment-anything-with-output_64_1.png | 3 - .../237-segment-anything-with-output_68_1.jpg | 3 + .../237-segment-anything-with-output_68_1.png | 3 + .../237-segment-anything-with-output_76_0.png | 3 - .../237-segment-anything-with-output_78_1.png | 3 - .../237-segment-anything-with-output_80_0.png | 3 + .../237-segment-anything-with-output_82_1.jpg | 3 + .../237-segment-anything-with-output_82_1.png | 3 + .../index.html | 26 +- .../238-deep-floyd-if-with-output.rst | 224 ++- .../238-deep-floyd-if-with-output_27_3.png | 3 - .../238-deep-floyd-if-with-output_29_3.png | 4 +- .../238-deep-floyd-if-with-output_31_3.png | 3 + ...=> 238-deep-floyd-if-with-output_41_0.png} | 0 .../index.html | 10 +- ...=> 239-image-bind-convert-with-output.rst} | 209 ++- ...9-image-bind-convert-with-output_20_0.png} | 0 ...9-image-bind-convert-with-output_22_0.png} | 0 ...39-image-bind-convert-with-output_24_0.png | 3 + ...39-image-bind-convert-with-output_26_1.jpg | 3 + ...39-image-bind-convert-with-output_26_1.png | 3 + ...39-image-bind-convert-with-output_27_1.jpg | 3 + ...39-image-bind-convert-with-output_27_1.png | 3 + ...39-image-bind-convert-with-output_28_1.jpg | 3 + ...39-image-bind-convert-with-output_28_1.png | 3 + .../index.html | 15 + .../239-image-bind-with-output_20_0.png | 3 - .../239-image-bind-with-output_22_1.png | 3 - .../239-image-bind-with-output_23_1.png | 3 - .../239-image-bind-with-output_24_1.png | 3 - .../index.html | 12 - ...ly-2-instruction-following-with-output.rst | 238 ++- ...41-riffusion-text-to-music-with-output.rst | 305 +++- ...ffusion-text-to-music-with-output_13_0.png | 3 - ...ffusion-text-to-music-with-output_15_0.jpg | 3 + ...ffusion-text-to-music-with-output_15_0.png | 3 + .../index.html | 7 +- ...42-freevc-voice-conversion-with-output.rst | 208 ++- ...tflite-selfie-segmentation-with-output.rst | 117 +- ...-selfie-segmentation-with-output_25_0.png} | 0 ...e-selfie-segmentation-with-output_32_0.png | 3 - ...e-selfie-segmentation-with-output_33_0.png | 3 + .../index.html | 8 +- ...4-named-entity-recognition-with-output.rst | 508 ++++++ .../245-typo-detector-with-output.rst | 593 +++++++ ...6-depth-estimation-videpth-with-output.rst | 1027 ++++++++++++ ...th-estimation-videpth-with-output_48_2.png | 3 + ...th-estimation-videpth-with-output_53_2.png | 3 + .../index.html | 8 + .../247-code-language-id-with-output.rst | 683 ++++++++ .../248-stable-diffusion-xl-with-output.rst | 612 +++++++ ...8-stable-diffusion-xl-with-output_10_3.jpg | 3 + ...8-stable-diffusion-xl-with-output_10_3.png | 3 + ...8-stable-diffusion-xl-with-output_18_3.jpg | 3 + ...8-stable-diffusion-xl-with-output_18_3.png | 3 + ...8-stable-diffusion-xl-with-output_29_3.jpg | 3 + ...8-stable-diffusion-xl-with-output_29_3.png | 3 + .../index.html | 12 + ...249-oneformer-segmentation-with-output.rst | 418 +++++ ...neformer-segmentation-with-output_22_1.jpg | 3 + ...neformer-segmentation-with-output_22_1.png | 3 + ...neformer-segmentation-with-output_26_0.jpg | 3 + ...neformer-segmentation-with-output_26_0.png | 3 + .../index.html | 10 + ...low-training-openvino-nncf-with-output.rst | 221 +-- ...raining-openvino-nncf-with-output_2_15.png | 4 +- ...training-openvino-nncf-with-output_2_9.png | 4 +- .../index.html | 11 - ...nsorflow-training-openvino-with-output.rst | 333 ++-- ...low-training-openvino-with-output_56_1.png | 4 +- ...low-training-openvino-with-output_65_0.png | 4 +- ...ow-training-openvino-with-output_78_1.png} | 0 .../index.html | 28 +- ...uantization-aware-training-with-output.rst | 173 +- ...uantization-aware-training-with-output.rst | 166 +- ...zation-aware-training-with-output_6_1.png} | 0 .../index.html | 6 +- .../401-object-detection-with-output.rst | 176 +- .../401-object-detection-with-output_20_0.png | 3 - .../401-object-detection-with-output_21_0.png | 3 + .../index.html | 6 +- .../402-pose-estimation-with-output.rst | 133 +- .../402-pose-estimation-with-output_20_0.png | 3 - .../402-pose-estimation-with-output_21_0.png | 3 + .../index.html | 6 +- ...-action-recognition-webcam-with-output.rst | 159 +- ...on-recognition-webcam-with-output_19_0.png | 3 - ...on-recognition-webcam-with-output_21_0.png | 3 + .../index.html | 6 +- .../404-style-transfer-with-output.rst | 160 +- ...> 404-style-transfer-with-output_27_0.png} | 0 .../index.html | 6 +- .../405-paddle-ocr-webcam-with-output.rst | 155 +- ...405-paddle-ocr-webcam-with-output_30_0.png | 3 - ...405-paddle-ocr-webcam-with-output_32_0.png | 3 + .../index.html | 6 +- .../406-3D-pose-estimation-with-output.rst | 235 ++- .../407-person-tracking-with-output.rst | 195 ++- ... 407-person-tracking-with-output_17_3.png} | 0 .../407-person-tracking-with-output_23_0.png | 3 - .../407-person-tracking-with-output_27_0.png | 3 + .../index.html | 8 +- docs/notebooks/index.html | 312 ++-- docs/notebooks/notebook_utils-with-output.rst | 17 +- .../index.html | 18 +- .../notebook_utils-with-output_26_0.png | 4 +- docs/notebooks/notebooks_tags.json | 60 +- .../notebooks_with_binder_buttons.txt | 4 +- .../notebooks_with_colab_buttons.txt | 4 +- docs/tutorials.md | 159 +- 360 files changed, 19728 insertions(+), 5810 deletions(-) create mode 100644 docs/notebooks/108-gpu-device-with-output.rst create mode 100644 docs/notebooks/109-throughput-tricks-with-output.rst create mode 100644 docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_14_0.jpg create mode 100644 docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_17_0.jpg create mode 100644 docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_20_0.jpg create mode 100644 docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_22_0.jpg create mode 100644 docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_26_0.jpg create mode 100644 docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_28_0.jpg create mode 100644 docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_30_0.jpg create mode 100644 docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_33_0.png create mode 100644 docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_4_0.jpg create mode 100644 docs/notebooks/109-throughput-tricks-with-output_files/index.html rename docs/notebooks/111-yolov5-quantization-migration-with-output_files/{111-yolov5-quantization-migration-with-output_33_0.png => 111-yolov5-quantization-migration-with-output_34_0.png} (100%) rename docs/notebooks/111-yolov5-quantization-migration-with-output_files/{111-yolov5-quantization-migration-with-output_39_0.png => 111-yolov5-quantization-migration-with-output_40_0.png} (100%) rename docs/notebooks/113-image-classification-quantization-with-output_files/{113-image-classification-quantization-with-output_28_2.png => 113-image-classification-quantization-with-output_29_2.png} (100%) create mode 100644 docs/notebooks/115-async-api-with-output_files/115-async-api-with-output_21_0.png rename docs/notebooks/118-optimize-preprocessing-with-output_files/{118-optimize-preprocessing-with-output_12_1.png => 118-optimize-preprocessing-with-output_13_1.png} (100%) create mode 100644 docs/notebooks/121-convert-to-openvino-with-output.rst delete mode 100644 docs/notebooks/205-vision-background-removal-with-output_files/205-vision-background-removal-with-output_20_0.png create mode 100644 docs/notebooks/205-vision-background-removal-with-output_files/205-vision-background-removal-with-output_24_0.png rename docs/notebooks/206-vision-paddlegan-anime-with-output_files/{206-vision-paddlegan-anime-with-output_36_0.png => 206-vision-paddlegan-anime-with-output_37_0.png} (100%) rename docs/notebooks/207-vision-paddlegan-superresolution-with-output_files/{207-vision-paddlegan-superresolution-with-output_25_1.png => 207-vision-paddlegan-superresolution-with-output_27_1.png} (100%) rename docs/notebooks/207-vision-paddlegan-superresolution-with-output_files/{207-vision-paddlegan-superresolution-with-output_29_1.png => 207-vision-paddlegan-superresolution-with-output_31_1.png} (100%) rename docs/notebooks/207-vision-paddlegan-superresolution-with-output_files/{207-vision-paddlegan-superresolution-with-output_31_0.png => 207-vision-paddlegan-superresolution-with-output_33_0.png} (100%) rename docs/notebooks/208-optical-character-recognition-with-output_files/{208-optical-character-recognition-with-output_13_0.png => 208-optical-character-recognition-with-output_15_0.png} (100%) delete mode 100644 docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_23_0.png rename docs/notebooks/208-optical-character-recognition-with-output_files/{208-optical-character-recognition-with-output_25_0.jpg => 208-optical-character-recognition-with-output_27_0.jpg} (100%) create mode 100644 docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_0.png rename docs/notebooks/208-optical-character-recognition-with-output_files/{208-optical-character-recognition-with-output_25_10.jpg => 208-optical-character-recognition-with-output_27_10.jpg} (100%) rename docs/notebooks/208-optical-character-recognition-with-output_files/{208-optical-character-recognition-with-output_25_10.png => 208-optical-character-recognition-with-output_27_10.png} (100%) rename docs/notebooks/208-optical-character-recognition-with-output_files/{208-optical-character-recognition-with-output_25_2.jpg => 208-optical-character-recognition-with-output_27_2.jpg} (100%) rename docs/notebooks/208-optical-character-recognition-with-output_files/{208-optical-character-recognition-with-output_25_2.png => 208-optical-character-recognition-with-output_27_2.png} (100%) rename docs/notebooks/208-optical-character-recognition-with-output_files/{208-optical-character-recognition-with-output_25_4.jpg => 208-optical-character-recognition-with-output_27_4.jpg} (100%) rename docs/notebooks/208-optical-character-recognition-with-output_files/{208-optical-character-recognition-with-output_25_4.png => 208-optical-character-recognition-with-output_27_4.png} (100%) rename docs/notebooks/208-optical-character-recognition-with-output_files/{208-optical-character-recognition-with-output_25_6.jpg => 208-optical-character-recognition-with-output_27_6.jpg} (100%) rename docs/notebooks/208-optical-character-recognition-with-output_files/{208-optical-character-recognition-with-output_25_6.png => 208-optical-character-recognition-with-output_27_6.png} (100%) rename docs/notebooks/208-optical-character-recognition-with-output_files/{208-optical-character-recognition-with-output_25_8.jpg => 208-optical-character-recognition-with-output_27_8.jpg} (100%) rename docs/notebooks/208-optical-character-recognition-with-output_files/{208-optical-character-recognition-with-output_25_8.png => 208-optical-character-recognition-with-output_27_8.png} (100%) rename docs/notebooks/209-handwritten-ocr-with-output_files/{209-handwritten-ocr-with-output_20_0.png => 209-handwritten-ocr-with-output_21_0.png} (100%) rename docs/notebooks/209-handwritten-ocr-with-output_files/{209-handwritten-ocr-with-output_29_1.png => 209-handwritten-ocr-with-output_30_1.png} (100%) rename docs/notebooks/212-pyannote-speaker-diarization-with-output_files/{212-pyannote-speaker-diarization-with-output_25_0.png => 212-pyannote-speaker-diarization-with-output_27_0.png} (100%) delete mode 100644 docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_12_0.png create mode 100644 docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_18_0.png delete mode 100644 docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_20_0.png create mode 100644 docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_22_0.png rename docs/notebooks/216-attention-center-with-output_files/{216-attention-center-with-output_11_1.png => 216-attention-center-with-output_14_1.png} (100%) rename docs/notebooks/216-attention-center-with-output_files/{216-attention-center-with-output_13_1.png => 216-attention-center-with-output_16_1.png} (100%) rename docs/notebooks/217-vision-deblur-with-output_files/{217-vision-deblur-with-output_22_0.png => 217-vision-deblur-with-output_24_0.png} (100%) delete mode 100644 docs/notebooks/217-vision-deblur-with-output_files/217-vision-deblur-with-output_25_0.png create mode 100644 docs/notebooks/217-vision-deblur-with-output_files/217-vision-deblur-with-output_29_0.png rename docs/notebooks/218-vehicle-detection-and-recognition-with-output_files/{218-vehicle-detection-and-recognition-with-output_12_0.png => 218-vehicle-detection-and-recognition-with-output_13_0.png} (100%) rename docs/notebooks/218-vehicle-detection-and-recognition-with-output_files/{218-vehicle-detection-and-recognition-with-output_19_0.png => 218-vehicle-detection-and-recognition-with-output_20_0.png} (100%) rename docs/notebooks/218-vehicle-detection-and-recognition-with-output_files/{218-vehicle-detection-and-recognition-with-output_25_0.png => 218-vehicle-detection-and-recognition-with-output_26_0.png} (100%) create mode 100644 docs/notebooks/220-cross-lingual-books-alignment-with-output.rst create mode 100644 docs/notebooks/220-cross-lingual-books-alignment-with-output_files/220-cross-lingual-books-alignment-with-output_31_0.png create mode 100644 docs/notebooks/220-cross-lingual-books-alignment-with-output_files/220-cross-lingual-books-alignment-with-output_48_0.png create mode 100644 docs/notebooks/220-cross-lingual-books-alignment-with-output_files/index.html rename docs/notebooks/222-vision-image-colorization-with-output_files/{222-vision-image-colorization-with-output_18_0.png => 222-vision-image-colorization-with-output_20_0.png} (100%) rename docs/notebooks/222-vision-image-colorization-with-output_files/{222-vision-image-colorization-with-output_19_0.png => 222-vision-image-colorization-with-output_21_0.png} (100%) rename docs/notebooks/224-3D-segmentation-point-clouds-with-output_files/{224-3D-segmentation-point-clouds-with-output_13_0.png => 224-3D-segmentation-point-clouds-with-output_15_0.png} (100%) delete mode 100644 docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_29_1.png create mode 100644 docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_37_1.png rename docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/{225-stable-diffusion-text-to-image-with-output_35_1.png => 225-stable-diffusion-text-to-image-with-output_39_1.png} (100%) delete mode 100644 docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_22_0.jpg delete mode 100644 docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_22_0.png create mode 100644 docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_26_0.jpg create mode 100644 docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_26_0.png delete mode 100644 docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_38_0.jpg delete mode 100644 docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_38_0.png create mode 100644 docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_43_0.jpg create mode 100644 docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_43_0.png delete mode 100644 docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_8_0.jpg delete mode 100644 docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_8_0.png create mode 100644 docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_9_0.jpg create mode 100644 docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_9_0.png rename docs/notebooks/{228-clip-zero-shot-image-classification-with-output.rst => 228-clip-zero-shot-convert-with-output.rst} (62%) create mode 100644 docs/notebooks/228-clip-zero-shot-convert-with-output_files/228-clip-zero-shot-convert-with-output_12_0.png create mode 100644 docs/notebooks/228-clip-zero-shot-convert-with-output_files/228-clip-zero-shot-convert-with-output_17_0.png create mode 100644 docs/notebooks/228-clip-zero-shot-convert-with-output_files/228-clip-zero-shot-convert-with-output_4_0.png create mode 100644 docs/notebooks/228-clip-zero-shot-convert-with-output_files/index.html delete mode 100644 docs/notebooks/228-clip-zero-shot-image-classification-with-output_files/228-clip-zero-shot-image-classification-with-output_10_0.png delete mode 100644 docs/notebooks/228-clip-zero-shot-image-classification-with-output_files/228-clip-zero-shot-image-classification-with-output_15_0.png delete mode 100644 docs/notebooks/228-clip-zero-shot-image-classification-with-output_files/228-clip-zero-shot-image-classification-with-output_5_0.png delete mode 100644 docs/notebooks/228-clip-zero-shot-image-classification-with-output_files/index.html create mode 100644 docs/notebooks/228-clip-zero-shot-quantize-with-output.rst create mode 100644 docs/notebooks/228-clip-zero-shot-quantize-with-output_files/228-clip-zero-shot-quantize-with-output_16_0.png create mode 100644 docs/notebooks/228-clip-zero-shot-quantize-with-output_files/index.html rename docs/notebooks/230-yolov8-optimization-with-output_files/{230-yolov8-optimization-with-output_12_1.png => 230-yolov8-optimization-with-output_13_1.png} (100%) delete mode 100644 docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_14_1.png create mode 100644 docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_15_1.png rename docs/notebooks/230-yolov8-optimization-with-output_files/{230-yolov8-optimization-with-output_24_0.png => 230-yolov8-optimization-with-output_27_0.png} (100%) rename docs/notebooks/230-yolov8-optimization-with-output_files/{230-yolov8-optimization-with-output_26_0.png => 230-yolov8-optimization-with-output_29_0.png} (100%) delete mode 100644 docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_54_0.png delete mode 100644 docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_56_0.png create mode 100644 docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_59_0.png create mode 100644 docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_61_0.png delete mode 100644 docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_85_0.png create mode 100644 docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_91_0.png create mode 100644 docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_97_0.png create mode 100644 docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_99_0.png rename docs/notebooks/231-instruct-pix2pix-image-editing-with-output_files/{231-instruct-pix2pix-image-editing-with-output_23_0.png => 231-instruct-pix2pix-image-editing-with-output_25_0.png} (100%) delete mode 100644 docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_28_1.png create mode 100644 docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_31_1.png delete mode 100644 docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_34_1.png create mode 100644 docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_37_1.png delete mode 100644 docs/notebooks/232-clip-language-saliency-map-with-output_files/index.html rename docs/notebooks/233-blip-visual-language-processing-with-output_files/{233-blip-visual-language-processing-with-output_24_0.png => 233-blip-visual-language-processing-with-output_28_0.png} (100%) rename docs/notebooks/233-blip-visual-language-processing-with-output_files/{233-blip-visual-language-processing-with-output_26_0.png => 233-blip-visual-language-processing-with-output_30_0.png} (100%) rename docs/notebooks/233-blip-visual-language-processing-with-output_files/{233-blip-visual-language-processing-with-output_7_0.png => 233-blip-visual-language-processing-with-output_8_0.png} (100%) rename docs/notebooks/234-encodec-audio-compression-with-output_files/{234-encodec-audio-compression-with-output_36_1.png => 234-encodec-audio-compression-with-output_38_1.png} (100%) delete mode 100644 docs/notebooks/235-controlnet-stable-diffusion-with-output_files/235-controlnet-stable-diffusion-with-output_14_0.png create mode 100644 docs/notebooks/235-controlnet-stable-diffusion-with-output_files/235-controlnet-stable-diffusion-with-output_17_0.png rename docs/notebooks/236-stable-diffusion-v2-text-to-image-demo-with-output_files/{236-stable-diffusion-v2-text-to-image-demo-with-output_23_0.png => 236-stable-diffusion-v2-text-to-image-demo-with-output_25_0.png} (100%) delete mode 100644 docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_17_0.png create mode 100644 docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_21_0.png delete mode 100644 docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_24_0.png create mode 100644 docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_28_0.png delete mode 100644 docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_31_0.png create mode 100644 docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_39_0.png delete mode 100644 docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_40_0.png create mode 100644 docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_48_0.png delete mode 100644 docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_49_0.png create mode 100644 docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_53_0.png delete mode 100644 docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_64_1.png create mode 100644 docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_68_1.jpg create mode 100644 docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_68_1.png delete mode 100644 docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_76_0.png delete mode 100644 docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_78_1.png create mode 100644 docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_80_0.png create mode 100644 docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_82_1.jpg create mode 100644 docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_82_1.png delete mode 100644 docs/notebooks/238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_27_3.png create mode 100644 docs/notebooks/238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_31_3.png rename docs/notebooks/238-deep-floyd-if-with-output_files/{238-deep-floyd-if-with-output_39_0.png => 238-deep-floyd-if-with-output_41_0.png} (100%) rename docs/notebooks/{239-image-bind-with-output.rst => 239-image-bind-convert-with-output.rst} (99%) rename docs/notebooks/{239-image-bind-with-output_files/239-image-bind-with-output_16_0.png => 239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_20_0.png} (100%) rename docs/notebooks/{239-image-bind-with-output_files/239-image-bind-with-output_18_0.png => 239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_22_0.png} (100%) create mode 100644 docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_24_0.png create mode 100644 docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_26_1.jpg create mode 100644 docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_26_1.png create mode 100644 docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_27_1.jpg create mode 100644 docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_27_1.png create mode 100644 docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_28_1.jpg create mode 100644 docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_28_1.png create mode 100644 docs/notebooks/239-image-bind-convert-with-output_files/index.html delete mode 100644 docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_20_0.png delete mode 100644 docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_22_1.png delete mode 100644 docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_23_1.png delete mode 100644 docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_24_1.png delete mode 100644 docs/notebooks/239-image-bind-with-output_files/index.html delete mode 100644 docs/notebooks/241-riffusion-text-to-music-with-output_files/241-riffusion-text-to-music-with-output_13_0.png create mode 100644 docs/notebooks/241-riffusion-text-to-music-with-output_files/241-riffusion-text-to-music-with-output_15_0.jpg create mode 100644 docs/notebooks/241-riffusion-text-to-music-with-output_files/241-riffusion-text-to-music-with-output_15_0.png rename docs/notebooks/243-tflite-selfie-segmentation-with-output_files/{243-tflite-selfie-segmentation-with-output_24_0.png => 243-tflite-selfie-segmentation-with-output_25_0.png} (100%) delete mode 100644 docs/notebooks/243-tflite-selfie-segmentation-with-output_files/243-tflite-selfie-segmentation-with-output_32_0.png create mode 100644 docs/notebooks/243-tflite-selfie-segmentation-with-output_files/243-tflite-selfie-segmentation-with-output_33_0.png create mode 100644 docs/notebooks/244-named-entity-recognition-with-output.rst create mode 100644 docs/notebooks/245-typo-detector-with-output.rst create mode 100644 docs/notebooks/246-depth-estimation-videpth-with-output.rst create mode 100644 docs/notebooks/246-depth-estimation-videpth-with-output_files/246-depth-estimation-videpth-with-output_48_2.png create mode 100644 docs/notebooks/246-depth-estimation-videpth-with-output_files/246-depth-estimation-videpth-with-output_53_2.png create mode 100644 docs/notebooks/246-depth-estimation-videpth-with-output_files/index.html create mode 100644 docs/notebooks/247-code-language-id-with-output.rst create mode 100644 docs/notebooks/248-stable-diffusion-xl-with-output.rst create mode 100644 docs/notebooks/248-stable-diffusion-xl-with-output_files/248-stable-diffusion-xl-with-output_10_3.jpg create mode 100644 docs/notebooks/248-stable-diffusion-xl-with-output_files/248-stable-diffusion-xl-with-output_10_3.png create mode 100644 docs/notebooks/248-stable-diffusion-xl-with-output_files/248-stable-diffusion-xl-with-output_18_3.jpg create mode 100644 docs/notebooks/248-stable-diffusion-xl-with-output_files/248-stable-diffusion-xl-with-output_18_3.png create mode 100644 docs/notebooks/248-stable-diffusion-xl-with-output_files/248-stable-diffusion-xl-with-output_29_3.jpg create mode 100644 docs/notebooks/248-stable-diffusion-xl-with-output_files/248-stable-diffusion-xl-with-output_29_3.png create mode 100644 docs/notebooks/248-stable-diffusion-xl-with-output_files/index.html create mode 100644 docs/notebooks/249-oneformer-segmentation-with-output.rst create mode 100644 docs/notebooks/249-oneformer-segmentation-with-output_files/249-oneformer-segmentation-with-output_22_1.jpg create mode 100644 docs/notebooks/249-oneformer-segmentation-with-output_files/249-oneformer-segmentation-with-output_22_1.png create mode 100644 docs/notebooks/249-oneformer-segmentation-with-output_files/249-oneformer-segmentation-with-output_26_0.jpg create mode 100644 docs/notebooks/249-oneformer-segmentation-with-output_files/249-oneformer-segmentation-with-output_26_0.png create mode 100644 docs/notebooks/249-oneformer-segmentation-with-output_files/index.html delete mode 100644 docs/notebooks/301-tensorflow-training-openvino-nncf-with-output_files/index.html rename docs/notebooks/301-tensorflow-training-openvino-with-output_files/{301-tensorflow-training-openvino-with-output_77_1.png => 301-tensorflow-training-openvino-with-output_78_1.png} (100%) rename docs/notebooks/305-tensorflow-quantization-aware-training-with-output_files/{305-tensorflow-quantization-aware-training-with-output_5_1.png => 305-tensorflow-quantization-aware-training-with-output_6_1.png} (100%) delete mode 100644 docs/notebooks/401-object-detection-with-output_files/401-object-detection-with-output_20_0.png create mode 100644 docs/notebooks/401-object-detection-with-output_files/401-object-detection-with-output_21_0.png delete mode 100644 docs/notebooks/402-pose-estimation-with-output_files/402-pose-estimation-with-output_20_0.png create mode 100644 docs/notebooks/402-pose-estimation-with-output_files/402-pose-estimation-with-output_21_0.png delete mode 100644 docs/notebooks/403-action-recognition-webcam-with-output_files/403-action-recognition-webcam-with-output_19_0.png create mode 100644 docs/notebooks/403-action-recognition-webcam-with-output_files/403-action-recognition-webcam-with-output_21_0.png rename docs/notebooks/404-style-transfer-with-output_files/{404-style-transfer-with-output_25_0.png => 404-style-transfer-with-output_27_0.png} (100%) delete mode 100644 docs/notebooks/405-paddle-ocr-webcam-with-output_files/405-paddle-ocr-webcam-with-output_30_0.png create mode 100644 docs/notebooks/405-paddle-ocr-webcam-with-output_files/405-paddle-ocr-webcam-with-output_32_0.png rename docs/notebooks/407-person-tracking-with-output_files/{407-person-tracking-with-output_13_3.png => 407-person-tracking-with-output_17_3.png} (100%) delete mode 100644 docs/notebooks/407-person-tracking-with-output_files/407-person-tracking-with-output_23_0.png create mode 100644 docs/notebooks/407-person-tracking-with-output_files/407-person-tracking-with-output_27_0.png diff --git a/docs/nbdoc/consts.py b/docs/nbdoc/consts.py index e8e7edd287b..bc8d7dc08c1 100644 --- a/docs/nbdoc/consts.py +++ b/docs/nbdoc/consts.py @@ -8,7 +8,7 @@ repo_owner = "openvinotoolkit" repo_name = "openvino_notebooks" -artifacts_link = "http://repository.toolbox.iotg.sclab.intel.com/projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/" +artifacts_link = "http://repository.toolbox.iotg.sclab.intel.com/projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/" blacklisted_extensions = ['.xml', '.bin'] diff --git a/docs/notebooks/001-hello-world-with-output.rst b/docs/notebooks/001-hello-world-with-output.rst index c8b6a714ae7..1b8752b43cf 100644 --- a/docs/notebooks/001-hello-world-with-output.rst +++ b/docs/notebooks/001-hello-world-with-output.rst @@ -1,6 +1,8 @@ Hello Image Classification ========================== +.. _top: + This basic introduction to OpenVINO™ shows how to do inference with an image classification model. @@ -11,10 +13,19 @@ Zoo `__ is used in this tutorial. For more information about how OpenVINO IR models are created, refer to the `TensorFlow to OpenVINO <101-tensorflow-classification-to-openvino-with-output.html>`__ -tutorial. +tutorial. -Imports -------- +**Table of contents**: + +- `Imports <#imports>`__ +- `Download the Model and data samples <#download-the-model-and-data-samples>`__ +- `Select inference device <#select-inference-device>`__ +- `Load the Model <#load-the-model>`__ +- `Load an Image <#load-an-image>`__ +- `Do Inference <#do-inference>`__ + +Imports `⇑ <#top>`__ +############################################ .. code:: ipython3 @@ -29,8 +40,8 @@ Imports sys.path.append("../utils") from notebook_utils import download_file -Download the Model and data samples ------------------------------------ +Download the Model and data samples `⇑ <#top>`__ +######################################################################## .. code:: ipython3 @@ -63,10 +74,10 @@ Download the Model and data samples artifacts/v3-small_224_1.0_float.bin: 0%| | 0.00/4.84M [00:00`__ +############################################################ -select device from dropdown list for running inference using OpenVINO +Select device from dropdown list for running inference using OpenVINO: .. code:: ipython3 @@ -91,8 +102,8 @@ select device from dropdown list for running inference using OpenVINO -Load the Model --------------- +Load the Model `⇑ <#top>`__ +################################################### .. code:: ipython3 @@ -102,8 +113,8 @@ Load the Model output_layer = compiled_model.output(0) -Load an Image -------------- +Load an Image `⇑ <#top>`__ +################################################## .. code:: ipython3 @@ -122,8 +133,8 @@ Load an Image .. image:: 001-hello-world-with-output_files/001-hello-world-with-output_10_0.png -Do Inference ------------- +Do Inference `⇑ <#top>`__ +################################################# .. code:: ipython3 diff --git a/docs/notebooks/001-hello-world-with-output_files/index.html b/docs/notebooks/001-hello-world-with-output_files/index.html index 07011e32ec2..790c2604103 100644 --- a/docs/notebooks/001-hello-world-with-output_files/index.html +++ b/docs/notebooks/001-hello-world-with-output_files/index.html @@ -1,7 +1,7 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/001-hello-world-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/001-hello-world-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/001-hello-world-with-output_files/


../
-001-hello-world-with-output_10_0.png               12-Jul-2023 00:11              387941
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/001-hello-world-with-output_files/


../
+001-hello-world-with-output_10_0.png               16-Aug-2023 01:31              387941
 

diff --git a/docs/notebooks/002-openvino-api-with-output.rst b/docs/notebooks/002-openvino-api-with-output.rst index 64332fbd0c3..1c2d0119965 100644 --- a/docs/notebooks/002-openvino-api-with-output.rst +++ b/docs/notebooks/002-openvino-api-with-output.rst @@ -4,29 +4,27 @@ OpenVINO™ Runtime API Tutorial This notebook explains the basics of the OpenVINO Runtime API. It covers: -- `Loading OpenVINO Runtime and Showing - Info <#Loading-OpenVINO-Runtime-and-Showing-Info>`__ -- `Loading a Model <#Loading-a-Model>`__ +- `Loading OpenVINO Runtime and Showing Info <#loading-openvino-runtime-and-showing-info>`__ +- `Loading a Model <#loading-a-model>`__ - - `OpenVINO IR Model <#OpenVINO-IR-Model>`__ - - `ONNX Model <#ONNX-Model>`__ - - `PaddlePaddle Model <#PaddlePaddle-Model>`__ - - `TensorFlow Model <#TensorFlow-Model>`__ - - `TensorFlow Lite Model <#TensorFlow-Lite-Model>`__ + - `OpenVINO IR Model <#openvino-ir-model>`__ + - `ONNX Model <#onnx-model>`__ + - `PaddlePaddle Model <#paddlepaddle-model>`__ + - `TensorFlow Model <#tensorflow-model>`__ + - `TensorFlow Lite Model <#tensorflow-lite-model>`__ -- `Getting Information about a - Model <#Getting-Information-about-a-Model>`__ +- `Getting Information about a Model <#getting-information-about-a-model>`__ - - `Model Inputs <#Model-Inputs>`__ - - `Model Outputs <#Model-Outputs>`__ + - `Model Inputs <#model-inputs>`__ + - `Model Outputs <#model-outputs>`__ -- `Doing Inference on a Model <#Doing-Inference-on-a-Model>`__ -- `Reshaping and Resizing <#Reshaping-and-Resizing>`__ +- `Doing Inference on a Model <#doing-inference-on-a-model>`__ +- `Reshaping and Resizing <#reshaping-and-resizing>`__ - - `Change Image Size <#Change-Image-Size>`__ - - `Change Batch Size <#Change-Batch-Size>`__ + - `Change Image Size <#change-image-size>`__ + - `Change Batch Size <#change-batch-size>`__ -- `Caching a Model <#Caching-a-Model>`__ +- `Caching a Model <#caching-a-model>`__ The notebook is divided into sections with headers. The next cell contains global requirements installation and imports. Each section is @@ -54,12 +52,12 @@ same. .. parsed-literal:: - Requirement already satisfied: requests in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (2.31.0) - Requirement already satisfied: tqdm in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (4.65.0) - Requirement already satisfied: charset-normalizer<4,>=2 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests) (3.2.0) - Requirement already satisfied: idna<4,>=2.5 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests) (3.4) - Requirement already satisfied: urllib3<3,>=1.21.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests) (1.26.16) - Requirement already satisfied: certifi>=2017.4.17 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests) (2023.5.7) + Requirement already satisfied: requests in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (2.31.0) + Requirement already satisfied: tqdm in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (4.66.1) + Requirement already satisfied: charset-normalizer<4,>=2 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests) (3.2.0) + Requirement already satisfied: idna<4,>=2.5 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests) (3.4) + Requirement already satisfied: urllib3<3,>=1.21.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests) (1.26.16) + Requirement already satisfied: certifi>=2017.4.17 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests) (2023.7.22) Loading OpenVINO Runtime and Showing Info @@ -116,7 +114,7 @@ OpenVINO IR Model An OpenVINO IR (Intermediate Representation) model consists of an ``.xml`` file, containing information about network topology, and a ``.bin`` file, containing the weights and biases binary data. Models in -OpenVINO IR format are obtained by using Model Optimizer tool. The +OpenVINO IR format are obtained by using model conversion API. The ``read_model()`` function expects the ``.bin`` weights file to have the same filename and be located in the same directory as the ``.xml`` file: ``model_weights_file == Path(model_xml).with_suffix(".bin")``. If this @@ -124,16 +122,16 @@ is the case, specifying the weights file is optional. If the weights file has a different filename, it can be specified using the ``weights`` parameter in ``read_model()``. -The OpenVINO `Model -Optimizer `__ -tool is used to convert models to OpenVINO IR format. Model Optimizer -reads the original model and creates an OpenVINO IR model (.xml and .bin -files) so inference can be performed without delays due to format -conversion. Optionally, Model Optimizer can adjust the model to be more -suitable for inference, for example, by alternating input shapes, -embedding preprocessing and cutting training parts off. For information -on how to convert your existing TensorFlow, PyTorch or ONNX model to -OpenVINO IR format with Model Optimizer, refer to the +The OpenVINO `model conversion +API `__ +tool is used to convert models to OpenVINO IR format. Model conversion +API reads the original model and creates an OpenVINO IR model (``.xml`` +and ``.bin`` files) so inference can be performed without delays due to +format conversion. Optionally, model conversion API can adjust the model +to be more suitable for inference, for example, by alternating input +shapes, embedding preprocessing and cutting training parts off. For +information on how to convert your existing TensorFlow, PyTorch or ONNX +model to OpenVINO IR format with model conversion API, refer to the `tensorflow-to-openvino <101-tensorflow-classification-to-openvino-with-output.html>`__ and `pytorch-onnx-to-openvino <102-pytorch-onnx-to-openvino-with-output.html>`__ @@ -165,7 +163,7 @@ notebooks. .. parsed-literal:: - PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/002-openvino-api/model/classification.bin') + PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/002-openvino-api/model/classification.bin') @@ -212,7 +210,7 @@ points to the filename of an ONNX model. .. parsed-literal:: - PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/002-openvino-api/model/segmentation.onnx') + PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/002-openvino-api/model/segmentation.onnx') @@ -268,7 +266,7 @@ without any conversion step. Pass the filename with extension to .. parsed-literal:: - PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/002-openvino-api/model/inference.pdiparams') + PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/002-openvino-api/model/inference.pdiparams') @@ -292,13 +290,15 @@ TensorFlow Model ~~~~~~~~~~~~~~~~ TensorFlow models saved in frozen graph format can also be passed to -``read_model`` starting in OpenVINO 2022.3. - -.. note:: - - * Directly loading TensorFlow models is available as a preview feature in the OpenVINO 2022.3 release. Fully functional support will be provided in the upcoming 2023 releases. - * Currently support is limited to only frozen graph inference format. Other TensorFlow model formats must be converted to OpenVINO IR using `Model Optimizer `__. +``read_model`` starting in OpenVINO 2022.3. + **NOTE**: Directly loading TensorFlow models is available as a + preview feature in the OpenVINO 2022.3 release. Fully functional + support will be provided in the upcoming 2023 releases. Currently + support is limited to only frozen graph inference format. Other + TensorFlow model formats must be converted to OpenVINO IR using + `model conversion + API `__. .. code:: ipython3 @@ -318,7 +318,7 @@ TensorFlow models saved in frozen graph format can also be passed to .. parsed-literal:: - PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/002-openvino-api/model/classification.pb') + PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/002-openvino-api/model/classification.pb') @@ -370,7 +370,7 @@ It is pre-trained model optimized to work with TensorFlow Lite. .. parsed-literal:: - PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/002-openvino-api/model/classification.tflite') + PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/002-openvino-api/model/classification.tflite') @@ -395,7 +395,7 @@ Getting Information about a Model The OpenVINO Model instance stores information about the model. Information about the inputs and outputs of the model are in ``model.inputs`` and ``model.outputs``. These are also properties of the -CompiledModel instance. While using ``model.inputs`` and +``CompiledModel`` instance. While using ``model.inputs`` and ``model.outputs`` in the cells below, you can also use ``compiled_model.inputs`` and ``compiled_model.outputs``. @@ -419,7 +419,7 @@ CompiledModel instance. While using ``model.inputs`` and .. parsed-literal:: - PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/002-openvino-api/model/classification.bin') + PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/002-openvino-api/model/classification.bin') @@ -581,8 +581,8 @@ on a model, first create an inference request by calling the ``compiled_model`` that was loaded with ``compile_model()``. Then, call the ``infer()`` method of ``InferRequest``. It expects one argument: ``inputs``. This is a dictionary that maps input layer names to input -data or list of input data in np.ndarray format, where the position of -the input tensor corresponds to input index. If a model has a single +data or list of input data in ``np.ndarray`` format, where the position +of the input tensor corresponds to input index. If a model has a single input, wrapping to a dictionary or list can be omitted. .. code:: ipython3 @@ -612,7 +612,7 @@ input, wrapping to a dictionary or list can be omitted. .. parsed-literal:: - PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/002-openvino-api/model/classification.bin') + PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/002-openvino-api/model/classification.bin') @@ -707,10 +707,10 @@ add the ``N`` dimension (where ``N``\ = 1) by calling the **Do inference** Now that the input data is in the right shape, run inference. The -CompiledModel inference result is a dictionary where keys are the Output -class instances (the same keys in ``compiled_model.outputs`` that can -also be obtained with ``compiled_model.output(index)``) and values - -predicted result in np.array format. +``CompiledModel`` inference result is a dictionary where keys are the +Output class instances (the same keys in ``compiled_model.outputs`` that +can also be obtained with ``compiled_model.output(index)``) and values - +predicted result in ``np.array`` format. .. code:: ipython3 @@ -797,7 +797,7 @@ input shape. .. parsed-literal:: - PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/002-openvino-api/model/segmentation.bin') + PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/002-openvino-api/model/segmentation.bin') @@ -948,7 +948,7 @@ the cache. .. parsed-literal:: - PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/002-openvino-api/model/classification.bin') + PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/002-openvino-api/model/classification.bin') diff --git a/docs/notebooks/003-hello-segmentation-with-output.rst b/docs/notebooks/003-hello-segmentation-with-output.rst index 07bb396be0c..664854ae8d3 100644 --- a/docs/notebooks/003-hello-segmentation-with-output.rst +++ b/docs/notebooks/003-hello-segmentation-with-output.rst @@ -1,6 +1,8 @@ Hello Image Segmentation ======================== +.. _top: + A very basic introduction to using segmentation models with OpenVINO™. In this tutorial, a pre-trained @@ -10,8 +12,19 @@ Zoo `__ is used. ADAS stands for Advanced Driver Assistance Services. The model recognizes four classes: background, road, curb and mark. -Imports -------- +**Table of contents**: + +- `Imports <#imports>`__ +- `Download model weights <#download-model-weights>`__ +- `Select inference device <#select-inference-device>`__ +- `Load the Model <#load-the-model>`__ +- `Load an Image <#load-an-image>`__ +- `Do Inference <#do-inference>`__ +- `Prepare Data for Visualization <#prepare-data-for-visualization>`__ +- `Visualize data <#visualize-data>`__ + +Imports `⇑ <#top>`__ +######################################### .. code:: ipython3 @@ -24,8 +37,9 @@ Imports sys.path.append("../utils") from notebook_utils import segmentation_map_to_image, download_file -Download model weights ----------------------- +Download model weights `⇑ <#top>`__ +############################################################################################################################# + .. code:: ipython3 @@ -61,10 +75,11 @@ Download model weights model/road-segmentation-adas-0001.bin: 0%| | 0.00/720k [00:00`__ +############################################################################################################################# -select device from dropdown list for running inference using OpenVINO + +Select device from dropdown list for running inference using OpenVINO: .. code:: ipython3 @@ -89,8 +104,9 @@ select device from dropdown list for running inference using OpenVINO -Load the Model --------------- +Load the Model `⇑ <#top>`__ +############################################################################################################################# + .. code:: ipython3 @@ -102,11 +118,10 @@ Load the Model input_layer_ir = compiled_model.input(0) output_layer_ir = compiled_model.output(0) -Load an Image -------------- +Load an Image `⇑ <#top>`__ +############################################################################################################################# -A sample image from the `Mapillary -Vistas `__ dataset is +A sample image from the `Mapillary Vistas `__ dataset is provided. .. code:: ipython3 @@ -134,7 +149,7 @@ provided. .. parsed-literal:: - + @@ -142,8 +157,9 @@ provided. .. image:: 003-hello-segmentation-with-output_files/003-hello-segmentation-with-output_10_1.png -Do Inference ------------- +Do Inference `⇑ <#top>`__ +############################################################################################################################# + .. code:: ipython3 @@ -159,7 +175,7 @@ Do Inference .. parsed-literal:: - + @@ -167,8 +183,9 @@ Do Inference .. image:: 003-hello-segmentation-with-output_files/003-hello-segmentation-with-output_12_1.png -Prepare Data for Visualization ------------------------------- +Prepare Data for Visualization `⇑ <#top>`__ +############################################################################################################################# + .. code:: ipython3 @@ -185,8 +202,9 @@ Prepare Data for Visualization # Create an image with mask. image_with_mask = cv2.addWeighted(resized_mask, alpha, rgb_image, 1 - alpha, 0) -Visualize data --------------- +Visualize data `⇑ <#top>`__ +############################################################################################################################# + .. code:: ipython3 diff --git a/docs/notebooks/003-hello-segmentation-with-output_files/index.html b/docs/notebooks/003-hello-segmentation-with-output_files/index.html index 47b8c0e7163..2502da6e407 100644 --- a/docs/notebooks/003-hello-segmentation-with-output_files/index.html +++ b/docs/notebooks/003-hello-segmentation-with-output_files/index.html @@ -1,9 +1,9 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/003-hello-segmentation-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/003-hello-segmentation-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/003-hello-segmentation-with-output_files/


../
-003-hello-segmentation-with-output_10_1.png        12-Jul-2023 00:11              249032
-003-hello-segmentation-with-output_12_1.png        12-Jul-2023 00:11               20550
-003-hello-segmentation-with-output_16_0.png        12-Jul-2023 00:11              260045
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/003-hello-segmentation-with-output_files/


../
+003-hello-segmentation-with-output_10_1.png        16-Aug-2023 01:31              249032
+003-hello-segmentation-with-output_12_1.png        16-Aug-2023 01:31               20550
+003-hello-segmentation-with-output_16_0.png        16-Aug-2023 01:31              260045
 

diff --git a/docs/notebooks/004-hello-detection-with-output.rst b/docs/notebooks/004-hello-detection-with-output.rst index d15393a8548..35d47be09d1 100644 --- a/docs/notebooks/004-hello-detection-with-output.rst +++ b/docs/notebooks/004-hello-detection-with-output.rst @@ -1,6 +1,8 @@ Hello Object Detection ====================== +.. _top: + A very basic introduction to using object detection models with OpenVINO™. @@ -14,10 +16,20 @@ shape of ``[100, 5]``. Each detected text box is stored in the ``(x_min, y_min)`` are the coordinates of the top left bounding box corner, ``(x_max, y_max)`` are the coordinates of the bottom right bounding box corner and ``conf`` is the confidence for the predicted -class. +class. -Imports -------- +**Table of contents**: + +- `Imports <#imports>`__ +- `Download model weights <#download-model-weights>`__ +- `Select inference device <#select-inference-device>`__ +- `Load the Model <#load-the-model>`__ +- `Load an Image <#load-an-image>`__ +- `Do Inference <#do-inference>`__ +- `Visualize Results <#visualize-results>`__ + +Imports `⇑ <#top>`__ +######################################## .. code:: ipython3 @@ -31,8 +43,8 @@ Imports sys.path.append("../utils") from notebook_utils import download_file -Download model weights ----------------------- +Download model weights `⇑ <#top>`__ +####################################################### .. code:: ipython3 @@ -67,10 +79,10 @@ Download model weights model/horizontal-text-detection-0001.bin: 0%| | 0.00/7.39M [00:00`__ +########################################################### -select device from dropdown list for running inference using OpenVINO +Select device from dropdown list for running inference using OpenVINO: .. code:: ipython3 @@ -95,8 +107,8 @@ select device from dropdown list for running inference using OpenVINO -Load the Model --------------- +Load the Model `⇑ <#top>`__ +############################################### .. code:: ipython3 @@ -108,8 +120,8 @@ Load the Model input_layer_ir = compiled_model.input(0) output_layer_ir = compiled_model.output("boxes") -Load an Image -------------- +Load an Image `⇑ <#top>`__ +############################################## .. code:: ipython3 @@ -132,8 +144,8 @@ Load an Image .. image:: 004-hello-detection-with-output_files/004-hello-detection-with-output_10_0.png -Do Inference ------------- +Do Inference `⇑ <#top>`__ +############################################## .. code:: ipython3 @@ -143,8 +155,8 @@ Do Inference # Remove zero only boxes. boxes = boxes[~np.all(boxes == 0, axis=1)] -Visualize Results ------------------ +Visualize Results `⇑ <#top>`__ +################################################## .. code:: ipython3 diff --git a/docs/notebooks/004-hello-detection-with-output_files/index.html b/docs/notebooks/004-hello-detection-with-output_files/index.html index 2915430f935..625caac9573 100644 --- a/docs/notebooks/004-hello-detection-with-output_files/index.html +++ b/docs/notebooks/004-hello-detection-with-output_files/index.html @@ -1,8 +1,8 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/004-hello-detection-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/004-hello-detection-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/004-hello-detection-with-output_files/


../
-004-hello-detection-with-output_10_0.png           12-Jul-2023 00:11              305482
-004-hello-detection-with-output_15_0.png           12-Jul-2023 00:11              457214
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/004-hello-detection-with-output_files/


../
+004-hello-detection-with-output_10_0.png           16-Aug-2023 01:31              305482
+004-hello-detection-with-output_15_0.png           16-Aug-2023 01:31              457214
 

diff --git a/docs/notebooks/101-tensorflow-classification-to-openvino-with-output.rst b/docs/notebooks/101-tensorflow-classification-to-openvino-with-output.rst index 5c60ccbb3a4..191f6d3e62e 100644 --- a/docs/notebooks/101-tensorflow-classification-to-openvino-with-output.rst +++ b/docs/notebooks/101-tensorflow-classification-to-openvino-with-output.rst @@ -1,18 +1,42 @@ Convert a TensorFlow Model to OpenVINO™ ======================================= -This short tutorial shows how to convert a TensorFlow -`MobileNetV3 `__ -image classification model to OpenVINO `Intermediate -Representation `__ -(OpenVINO IR) format, using `Model -Optimizer `__. -After creating the OpenVINO IR, load the model in `OpenVINO -Runtime `__ -and do inference with a sample image. +.. _top: + +| This short tutorial shows how to convert a TensorFlow + `MobileNetV3 `__ + image classification model to OpenVINO `Intermediate + Representation `__ + (OpenVINO IR) format, using `model conversion + API `__. + After creating the OpenVINO IR, load the model in `OpenVINO + Runtime `__ + and do inference with a sample image. + +| **Table of contents**: + +- `Imports <#imports>`__ +- `Settings <#settings>`__ +- `Download model <#download-model>`__ +- `Convert a Model to OpenVINO IR Format <#convert-a-model-to-openvino-ir-format>`__ + + - `Convert a TensorFlow Model to OpenVINO IR Format <#convert-a-tensorflow-model-to-openvino-ir-format>`__ + +- `Test Inference on the Converted Model <#test-inference-on-the-converted-model>`__ + + - `Load the Model <#load-the-model>`__ + +- `Select inference device <#select-inference-device>`__ + + - `Get Model Information <#get-model-information>`__ + - `Load an Image <#load-an-image>`__ + - `Do Inference <#do-inference>`__ + +- `Timing <#timing>`__ + +Imports `⇑ <#top>`__ +############################################################################################################################### -Imports -------- .. code:: ipython3 @@ -29,14 +53,15 @@ Imports .. parsed-literal:: - 2023-07-11 22:24:31.659014: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. - 2023-07-11 22:24:31.694466: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + 2023-08-15 22:26:34.199621: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-08-15 22:26:34.233464: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. - 2023-07-11 22:24:32.210417: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + 2023-08-15 22:26:34.746193: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT -Settings --------- +Settings `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -48,8 +73,9 @@ Settings ir_path = Path("model/v3-small_224_1.0_float.xml") -Download model --------------- +Download model `⇑ <#top>`__ +############################################################################################################################### + Load model using `tf.keras.applications api `__ @@ -68,7 +94,7 @@ and save it to the disk. .. parsed-literal:: - 2023-07-11 22:24:33.117393: W tensorflow/core/common_runtime/gpu/gpu_device.cc:1956] Cannot dlopen some GPU libraries. Please make sure the missing libraries mentioned above are installed properly if you would like to use GPU. Follow the guide at https://www.tensorflow.org/install/gpu for how to download and setup the required libraries for your platform. + 2023-08-15 22:26:35.659386: W tensorflow/core/common_runtime/gpu/gpu_device.cc:1956] Cannot dlopen some GPU libraries. Please make sure the missing libraries mentioned above are installed properly if you would like to use GPU. Follow the guide at https://www.tensorflow.org/install/gpu for how to download and setup the required libraries for your platform. Skipping registering GPU devices... @@ -79,9 +105,9 @@ and save it to the disk. .. parsed-literal:: - 2023-07-11 22:24:37.315661: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,1,1,1024] + 2023-08-15 22:26:39.846021: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,1,1,1024] [[{{node inputs}}]] - 2023-07-11 22:24:40.473876: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,1,1,1024] + 2023-08-15 22:26:42.992490: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,1,1,1024] [[{{node inputs}}]] WARNING:absl:Found untraced functions such as _jit_compiled_convolution_op, _jit_compiled_convolution_op, _jit_compiled_convolution_op, _jit_compiled_convolution_op, _jit_compiled_convolution_op while saving (showing 5 of 54). These functions will not be directly callable after loading. @@ -96,25 +122,27 @@ and save it to the disk. INFO:tensorflow:Assets written to: model/v3-small_224_1.0_float/assets -Convert a Model to OpenVINO IR Format -------------------------------------- +Convert a Model to OpenVINO IR Format `⇑ <#top>`__ +############################################################################################################################### -Convert a TensorFlow Model to OpenVINO IR Format -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ -Call the OpenVINO Model Optimizer Python API to convert the TensorFlow -model to OpenVINO IR. ``mo.convert_model`` function accept path to saved +Convert a TensorFlow Model to OpenVINO IR Format `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Use the model conversion Python API to convert the TensorFlow model to +OpenVINO IR. The ``mo.convert_model`` function accept path to saved model directory and returns OpenVINO Model class instance which -represents this model. Obtained model is ready to use and loading on -device using ``compile_model`` or can be saved on disk using -``serialize`` function. See the `Model Optimizer Developer -Guide `__ -for more information about Model Optimizer and TensorFlow models -conversion. +represents this model. Obtained model is ready to use and to be loaded +on a device using ``compile_model`` or can be saved on a disk using the +``serialize`` function. See the +`tutorial `__ +for more information about using model conversion API with TensorFlow +models. .. code:: ipython3 - # Run Model Optimizer if the IR model file does not exist + # Run model conversion API if the IR model file does not exist if not ir_path.exists(): print("Exporting TensorFlow model to IR... This may take a few minutes.") ov_model = mo.convert_model(saved_model_dir=model_path, input_shape=[[1, 224, 224, 3]], compress_to_fp16=True) @@ -128,21 +156,24 @@ conversion. Exporting TensorFlow model to IR... This may take a few minutes. -Test Inference on the Converted Model -------------------------------------- +Test Inference on the Converted Model `⇑ <#top>`__ +############################################################################################################################### + + +Load the Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Load the Model -~~~~~~~~~~~~~~ .. code:: ipython3 core = Core() model = core.read_model(ir_path) -Select inference device ------------------------ +Select inference device `⇑ <#top>`__ +############################################################################################################################### -select device from dropdown list for running inference using OpenVINO + +Select device from dropdown list for running inference using OpenVINO: .. code:: ipython3 @@ -170,8 +201,9 @@ select device from dropdown list for running inference using OpenVINO compiled_model = core.compile_model(model=model, device_name=device.value) -Get Model Information -~~~~~~~~~~~~~~~~~~~~~ +Get Model Information `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -179,8 +211,9 @@ Get Model Information output_key = compiled_model.output(0) network_input_shape = input_key.shape -Load an Image -~~~~~~~~~~~~~ +Load an Image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Load an image, resize it, and convert it to the input shape of the network. @@ -203,8 +236,9 @@ network. .. image:: 101-tensorflow-classification-to-openvino-with-output_files/101-tensorflow-classification-to-openvino-with-output_18_0.png -Do Inference -~~~~~~~~~~~~ +Do Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -228,8 +262,9 @@ Do Inference -Timing ------- +Timing `⇑ <#top>`__ +############################################################################################################################### + Measure the time it takes to do inference on thousand images. This gives an indication of performance. For more accurate benchmarking, use the @@ -258,5 +293,5 @@ performance. .. parsed-literal:: - IR model in OpenVINO Runtime/CPU: 0.0010 seconds per image, FPS: 998.88 + IR model in OpenVINO Runtime/CPU: 0.0010 seconds per image, FPS: 988.20 diff --git a/docs/notebooks/101-tensorflow-classification-to-openvino-with-output_files/index.html b/docs/notebooks/101-tensorflow-classification-to-openvino-with-output_files/index.html index 6c3bb2041d0..9b704c7bd1b 100644 --- a/docs/notebooks/101-tensorflow-classification-to-openvino-with-output_files/index.html +++ b/docs/notebooks/101-tensorflow-classification-to-openvino-with-output_files/index.html @@ -1,7 +1,7 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/101-tensorflow-classification-to-openvino-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/101-tensorflow-classification-to-openvino-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/101-tensorflow-classification-to-openvino-with-output_files/


../
-101-tensorflow-classification-to-openvino-with-..> 12-Jul-2023 00:11              387941
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/101-tensorflow-classification-to-openvino-with-output_files/


../
+101-tensorflow-classification-to-openvino-with-..> 16-Aug-2023 01:31              387941
 

diff --git a/docs/notebooks/102-pytorch-onnx-to-openvino-with-output.rst b/docs/notebooks/102-pytorch-onnx-to-openvino-with-output.rst index c407ff0749f..c310e83f56d 100644 --- a/docs/notebooks/102-pytorch-onnx-to-openvino-with-output.rst +++ b/docs/notebooks/102-pytorch-onnx-to-openvino-with-output.rst @@ -1,6 +1,8 @@ Convert a PyTorch Model to ONNX and OpenVINO™ IR ================================================ +.. _top: + This tutorial demonstrates step-by-step instructions on how to do inference on a PyTorch semantic segmentation model, using OpenVINO Runtime. @@ -28,16 +30,43 @@ all 80 classes, the segmentation model has been trained on 20 classes from the `PASCAL VOC `__ dataset: **background, aeroplane, bicycle, bird, boat, bottle, bus, car, cat, chair, cow, dining table, dog, horse, motorbike, person, potted -plant, sheep, sofa, train, tvmonitor** +plant, sheep, sofa, train, tv monitor** More information about the model is available in the `torchvision documentation `__ -Preparation ------------ +**Table of contents**: -Imports -~~~~~~~ +- `Preparation <#preparation>`__ + + - `Imports <#imports>`__ + - `Settings <#settings>`__ + - `Load Model <#load-model>`__ + +- `ONNX Model Conversion <#onnx-model-conversion>`__ + + - `Convert PyTorch model to ONNX <#convert-pytorch-model-to-onnx>`__ + - `Convert ONNX Model to OpenVINO IR Format <#convert-onnx-model-to-openvino-ir-format>`__ + +- `Show Results <#show-results>`__ + + - `Load and Preprocess an Input Image <#load-and-preprocess-an-input-image>`__ + - `Load the OpenVINO IR Network and Run Inference on the ONNX model <#load-the-openvino-ir-network-and-run-inference-on-the-onnx-model>`__ + + - `1. ONNX Model in OpenVINO Runtime <#onnx-model-in-openvino-runtime>`__ + - `Select an inference device <#select-an-inference-device>`__ + - `2. OpenVINO IR Model in OpenVINO Runtime <#openvino-ir-model-in-openvino-runtime>`__ + - `Select the inference device <#select-the-inference-device>`__ + +- `PyTorch Comparison <#pytorch-comparison>`__ +- `Performance Comparison <#performance-comparison>`__ +- `References <#references>`__ + +Preparation `⇑ <#top>`__ +######################################################################## + +Imports `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ .. code:: ipython3 @@ -55,8 +84,8 @@ Imports sys.path.append("../utils") from notebook_utils import segmentation_map_to_image, viz_result_image, SegmentationMap, Label, download_file -Settings -~~~~~~~~ +Settings `⇑ <#top>`__ +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ Set a name for the model, then define width and height of the image that will be used by the network during inference. According to the input @@ -77,16 +106,15 @@ transforms function, the model is pre-trained on images with a height of onnx_path.parent.mkdir() ir_path = onnx_path.with_suffix(".xml") -Load Model -~~~~~~~~~~ +Load Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ Generally, PyTorch models represent an instance of ``torch.nn.Module`` class, initialized by a state dictionary with model weights. Typical -steps for getting a pre-trained model: - -1. Create instance of model class +steps for getting a pre-trained model: 1. Create instance of model class 2. Load checkpoint state dict, which contains pre-trained model weights -3. Turn model to evaluation for switching some operations to inference mode +3. Turn model to evaluation for switching some operations to inference +mode The ``torchvision`` module provides a ready to use set of functions for model class initialization. We will use @@ -129,11 +157,11 @@ have not downloaded the model before. Loaded PyTorch LRASPP MobileNetV3 model -ONNX Model Conversion ---------------------- +ONNX Model Conversion `⇑ <#top>`__ +################################################################################ -Convert PyTorch model to ONNX -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Convert PyTorch model to ONNX `⇑ <#top>`__ +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ OpenVINO supports PyTorch models that are exported in ONNX format. We will use the ``torch.onnx.export`` function to obtain the ONNX model, @@ -172,14 +200,13 @@ line of the output will read: ONNX model exported to model/lraspp_mobilenet_v3_large.onnx. -Convert ONNX Model to OpenVINO IR Format -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Convert ONNX Model to OpenVINO IR Format `⇑ <#top>`__ +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Use Model Optimizer to convert the ONNX model to OpenVINO IR with -``FP16`` precision. The models are saved inside the current directory. -For more information about Model Optimizer, see the `Model Optimizer -Developer -Guide `__. +To convert the ONNX model to OpenVINO IR with ``FP16`` precision, use +model conversion API. The models are saved inside the current directory. +For more information on how to convert models, see this +`page `__. .. code:: ipython3 @@ -199,14 +226,14 @@ Guide `__ +###################################################################### Confirm that the segmentation results look as expected by comparing model predictions on the ONNX, OpenVINO IR and PyTorch models. -Load and Preprocess an Input Image -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Load and Preprocess an Input Image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ Images need to be normalized before propagating through the network. @@ -237,8 +264,8 @@ Images need to be normalized before propagating through the network. input_image = np.expand_dims(np.transpose(resized_image, (2, 0, 1)), 0) normalized_input_image = np.expand_dims(np.transpose(normalized_image, (2, 0, 1)), 0) -Load the OpenVINO IR Network and Run Inference on the ONNX model -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Load the OpenVINO IR Network and Run Inference on the ONNX model `⇑ <#top>`__ +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ OpenVINO Runtime can load ONNX models directly. First, load the ONNX model, do inference and show the results. Then, load the model that was @@ -246,8 +273,8 @@ converted to OpenVINO Intermediate Representation (OpenVINO IR) with Model Optimizer and do inference on that model, and show the results on an image. -1. ONNX Model in OpenVINO Runtime -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +1. ONNX Model in OpenVINO Runtime `⇑ <#top>`__ +------------------------------------------------------------------------------------------ .. code:: ipython3 @@ -257,10 +284,10 @@ an image. # Read model to OpenVINO Runtime model_onnx = core.read_model(model=onnx_path) -Select inference device -^^^^^^^^^^^^^^^^^^^^^^^ +Select an inference device `⇑ <#top>`__ +................................................................................... -select device from dropdown list for running inference using OpenVINO +Select a device from dropdown list for running inference using OpenVINO: .. code:: ipython3 @@ -339,13 +366,13 @@ be applied to each label for more convenient visualization. -2. OpenVINO IR Model in OpenVINO Runtime -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +2. OpenVINO IR Model in OpenVINO Runtime `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------ -Select inference device -^^^^^^^^^^^^^^^^^^^^^^^ +Select the inference device `⇑ <#top>`__ +..................................................................................... -select device from dropdown list for running inference using OpenVINO +Select a device from dropdown list for running inference using OpenVINO: .. code:: ipython3 @@ -389,8 +416,8 @@ select device from dropdown list for running inference using OpenVINO -PyTorch Comparison ------------------- +PyTorch Comparison `⇑ <#top>`__ +############################################################################ Do inference on the PyTorch model to verify that the output visually looks the same as the output on the ONNX/OpenVINO IR models. @@ -415,8 +442,8 @@ looks the same as the output on the ONNX/OpenVINO IR models. -Performance Comparison ----------------------- +Performance Comparison `⇑ <#top>`__ +################################################################################ Measure the time it takes to do inference on twenty images. This gives an indication of performance. For more accurate benchmarking, use the @@ -488,9 +515,9 @@ performance. .. parsed-literal:: - PyTorch model on CPU: 0.039 seconds per image, FPS: 25.80 - ONNX model in OpenVINO Runtime/CPU: 0.031 seconds per image, FPS: 31.95 - OpenVINO IR model in OpenVINO Runtime/CPU: 0.031 seconds per image, FPS: 32.67 + PyTorch model on CPU: 0.037 seconds per image, FPS: 27.19 + ONNX model in OpenVINO Runtime/CPU: 0.031 seconds per image, FPS: 32.33 + OpenVINO IR model in OpenVINO Runtime/CPU: 0.032 seconds per image, FPS: 31.72 **Show Device Information** @@ -508,8 +535,8 @@ performance. CPU: Intel(R) Core(TM) i9-10920X CPU @ 3.50GHz -References ----------- +References `⇑ <#top>`__ +###################################################################### - `Torchvision `__ - `Pytorch ONNX @@ -517,7 +544,7 @@ References - `PIP install openvino-dev `__ - `OpenVINO ONNX support `__ -- `Model Optimizer - Documentation `__ -- `Model Optimizer Pytorch conversion - guide `__ +- `Model Conversion API + documentation `__ +- `Converting Pytorch + model `__ diff --git a/docs/notebooks/102-pytorch-onnx-to-openvino-with-output_files/index.html b/docs/notebooks/102-pytorch-onnx-to-openvino-with-output_files/index.html index 12501257639..d28d774056a 100644 --- a/docs/notebooks/102-pytorch-onnx-to-openvino-with-output_files/index.html +++ b/docs/notebooks/102-pytorch-onnx-to-openvino-with-output_files/index.html @@ -1,9 +1,9 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/102-pytorch-onnx-to-openvino-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/102-pytorch-onnx-to-openvino-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/102-pytorch-onnx-to-openvino-with-output_files/


../
-102-pytorch-onnx-to-openvino-with-output_21_0.png  12-Jul-2023 00:11              465692
-102-pytorch-onnx-to-openvino-with-output_26_0.png  12-Jul-2023 00:11              465695
-102-pytorch-onnx-to-openvino-with-output_28_0.png  12-Jul-2023 00:11              465692
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/102-pytorch-onnx-to-openvino-with-output_files/


../
+102-pytorch-onnx-to-openvino-with-output_21_0.png  16-Aug-2023 01:31              465692
+102-pytorch-onnx-to-openvino-with-output_26_0.png  16-Aug-2023 01:31              465695
+102-pytorch-onnx-to-openvino-with-output_28_0.png  16-Aug-2023 01:31              465692
 

diff --git a/docs/notebooks/102-pytorch-to-openvino-with-output.rst b/docs/notebooks/102-pytorch-to-openvino-with-output.rst index 68fa9273364..310055a2ec8 100644 --- a/docs/notebooks/102-pytorch-to-openvino-with-output.rst +++ b/docs/notebooks/102-pytorch-to-openvino-with-output.rst @@ -1,6 +1,8 @@ Convert a PyTorch Model to OpenVINO™ IR ======================================= +.. _top: + This tutorial demonstrates step-by-step instructions on how to do inference on a PyTorch classification model using OpenVINO Runtime. Starting from OpenVINO 2023.0 release, OpenVINO supports direct PyTorch @@ -27,12 +29,45 @@ network design spaces that parametrize populations of networks. The overall process is analogous to the classic manual design of networks but elevated to the design space level. The RegNet design space provides simple and fast networks that work well across a wide range of flop -regimes. +regimes. -Prerequisites -------------- +**Table of contents**: -Install notebook dependecies +- `Prerequisites <#prerequisites>`__ +- `Load PyTorch Model <#load-pytorch-model>`__ + + - `Prepare Input Data <#prepare-input-data>`__ + - `Run PyTorch Model Inference <#run-pytorch-model-inference>`__ + - `Benchmark PyTorch Model Inference <#benchmark-pytorch-model-inference>`__ + +- `Convert PyTorch Model to OpenVINO Intermediate Representation <#convert-pytorch-model-to-openvino-intermediate-representation>`__ + + - `Select inference device <#select-inference-device>`__ + - `Run OpenVINO Model Inference <#run-openvino-model-inference>`__ + - `Benchmark OpenVINO Model Inference <#benchmark-openvino-model-inference>`__ + +- `Convert PyTorch Model with Static Input Shape <#convert-pytorch-model-with-static-input-shape>`__ + + - `Select inference device <#select-inference-device>`__ + - `Run OpenVINO Model Inference with Static Input Shape <#run-openvino-model-inference-with-static-input-shape>`__ + - `Benchmark OpenVINO Model Inference with Static Input Shape <#benchmark-openvino-model-inference-with-static-input-shape>`__ + +- `Convert TorchScript Model to OpenVINO Intermediate Representation <#convert-torchscript-model-to-openvino-intermediate-representation>`__ + + - `Scripted Model <#scripted-model>`__ + - `Benchmark Scripted Model Inference <#benchmark-scripted-model-inference>`__ + - `Convert PyTorch Scripted Model to OpenVINO Intermediate Representation <#convert-pytorch-scripted-model-to-openvino-intermediate-representation>`__ + - `Benchmark OpenVINO Model Inference Converted From Scripted Model <#benchmark-openvino-model-inference-converted-from-scripted-model>`__ + - `Traced Model <#traced-model>`__ + - `Benchmark Traced Model Inference <#benchmark-traced-model-inference>`__ + - `Convert PyTorch Traced Model to OpenVINO Intermediate Representation <#convert-pytorch-traced-model-to-openvino-intermediate-representation>`__ + - `Benchmark OpenVINO Model Inference Converted From Traced Model <#benchmark-openvino-model-inference-converted-from-traced-model>`__ + +Prerequisites `⇑ <#top>`__ +############################################################################################################################### + + +Install notebook dependencies .. code:: ipython3 @@ -64,8 +99,9 @@ Download input data and label map imagenet_classes = labels_file.open("r").read().splitlines() -Load PyTorch Model ------------------- +Load PyTorch Model `⇑ <#top>`__ +############################################################################################################################### + Generally, PyTorch models represent an instance of the ``torch.nn.Module`` class, initialized by a state dictionary with model @@ -73,7 +109,8 @@ weights. Typical steps for getting a pre-trained model: 1. Create an instance of a model class 2. Load checkpoint state dict, which contains pre-trained model weights -3. Turn the model to evaluation for switching some operations to inference mode +3. Turn the model to evaluation for switching some operations to + inference mode The ``torchvision`` module provides a ready-to-use set of functions for model class initialization. We will use @@ -94,8 +131,9 @@ enum ``RegNet_Y_800MF_Weights.DEFAULT``. # switch model to inference mode model.eval(); -Prepare Input Data -~~~~~~~~~~~~~~~~~~ +Prepare Input Data `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The code below demonstrates how to preprocess input data using a model-specific transforms module from ``torchvision``. After @@ -116,8 +154,9 @@ the first dimension. # Add batch dimension to image tensor input_tensor = img_transformed.unsqueeze(0) -Run PyTorch Model Inference -~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Run PyTorch Model Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The model returns a vector of probabilities in raw logits format, softmax can be applied to get normalized values in the [0, 1] range. For @@ -172,8 +211,9 @@ can be reused later. 5: hamper - 2.35% -Benchmark PyTorch Model Inference -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Benchmark PyTorch Model Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -185,11 +225,11 @@ Benchmark PyTorch Model Inference .. parsed-literal:: - 13.5 ms ± 3.76 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) + 13.2 ms ± 27.7 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) -Convert PyTorch Model to OpenVINO Intermediate Representation -------------------------------------------------------------- +Convert PyTorch Model to OpenVINO Intermediate Representation. `⇑ <#top>`__ +############################################################################################################################### Starting from the 2023.0 release OpenVINO supports direct PyTorch models conversion to OpenVINO Intermediate Representation (IR) format. Model @@ -211,13 +251,17 @@ device using ``core.compile_model`` or save on disk for next usage using ``openvino.runtime.serialize``. Optionally, we can provide additional parameters, such as: -* ``compress_to_fp16`` - flag to perform model weights compression into FP16 data format. It may reduce the required space for model storage on disk and give speedup for inference devices, where FP16 calculation is supported. -* ``example_input`` - input data sample which can be used for model tracing. -* ``input_shape`` - the shape of input tensor for conversion +- ``compress_to_fp16`` - flag to perform model weights compression into + FP16 data format. It may reduce the required space for model storage + on disk and give speedup for inference devices, where FP16 + calculation is supported. +- ``example_input`` - input data sample which can be used for model + tracing. +- ``input_shape`` - the shape of input tensor for conversion -and any other advanced options supported by Model Optimizer Python API. +and any other advanced options supported by model conversion Python API. More details can be found on this -`page `__ +`page `__ .. code:: ipython3 @@ -250,10 +294,11 @@ More details can be found on this -Select inference device -~~~~~~~~~~~~~~~~~~~~~~~ +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -select device from dropdown list for running inference using OpenVINO + +Select device from dropdown list for running inference using OpenVINO: .. code:: ipython3 @@ -298,8 +343,9 @@ select device from dropdown list for running inference using OpenVINO -Run OpenVINO Model Inference -~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Run OpenVINO Model Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -329,8 +375,9 @@ Run OpenVINO Model Inference 5: hamper - 2.35% -Benchmark OpenVINO Model Inference -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Benchmark OpenVINO Model Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -341,11 +388,12 @@ Benchmark OpenVINO Model Inference .. parsed-literal:: - 3.03 ms ± 5.57 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) + 3.03 ms ± 45.2 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) -Convert PyTorch Model with Static Input Shape ---------------------------------------------- +Convert PyTorch Model with Static Input Shape `⇑ <#top>`__ +############################################################################################################################### + The default conversion path preserves dynamic input shapes, in order if you want to convert the model with static shapes, you can explicitly @@ -377,10 +425,11 @@ reshaping example please check the following -Select inference device -~~~~~~~~~~~~~~~~~~~~~~~ +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -select device from dropdown list for running inference using OpenVINO + +Select device from dropdown list for running inference using OpenVINO: .. code:: ipython3 @@ -420,8 +469,9 @@ Now, we can see that input of our converted model is tensor of shape [1, 3, 224, 224] instead of [?, 3, ?, ?] reported by previously converted model. -Run OpenVINO Model Inference with Static Input Shape -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Run OpenVINO Model Inference with Static Input Shape `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -452,7 +502,7 @@ Run OpenVINO Model Inference with Static Input Shape Benchmark OpenVINO Model Inference with Static Input Shape -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +`⇑ <#top>`__ .. code:: ipython3 @@ -463,11 +513,11 @@ Benchmark OpenVINO Model Inference with Static Input Shape .. parsed-literal:: - 2.79 ms ± 26.2 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) + 2.77 ms ± 12.7 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) -Convert TorchScript Model to OpenVINO Intermediate Representation ------------------------------------------------------------------ +Convert TorchScript Model to OpenVINO Intermediate Representation. `⇑ <#top>`__ +############################################################################################################################### TorchScript is a way to create serializable and optimizable models from PyTorch code. Any TorchScript program can be saved from a Python process @@ -487,8 +537,9 @@ There are 2 possible ways to convert the PyTorch model to TorchScript: Let’s consider both approaches and their conversion into OpenVINO IR. -Scriped Model -~~~~~~~~~~~~~ +Scripted Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + ``torch.jit.script`` inspects model source code and compiles it to ``ScriptModule``. After compilation model can be used for inference or @@ -540,8 +591,9 @@ Reference `__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -552,14 +604,12 @@ Benchmark Scripted Model Inference .. parsed-literal:: - 12.7 ms ± 13.4 µs per loop (mean ± std. dev. of 7 runs, 10 loops each) + 12.6 ms ± 17.6 µs per loop (mean ± std. dev. of 7 runs, 10 loops each) -Convert PyTorch Scripted Model to OpenVINO Intermediate Representation -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ - -The conversion step for the scripted model to OpenVINO IR is similar to -the original PyTorch model. +Convert PyTorch Scripted Model to OpenVINO Intermediate +Representation `⇑ <#top>`__ The conversion step for the scripted model to +OpenVINO IR is similar to the original PyTorch model. .. code:: ipython3 @@ -596,7 +646,7 @@ the original PyTorch model. Benchmark OpenVINO Model Inference Converted From Scripted Model -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +`⇑ <#top>`__ .. code:: ipython3 @@ -607,11 +657,12 @@ Benchmark OpenVINO Model Inference Converted From Scripted Model .. parsed-literal:: - 3.1 ms ± 4.02 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) + 3.07 ms ± 5.58 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) -Traced Model -~~~~~~~~~~~~ +Traced Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Using ``torch.jit.trace``, you can turn an existing module or Python function into a TorchScript ``ScriptFunction`` or ``ScriptModule``. You @@ -619,10 +670,10 @@ must provide example inputs, and model will be executed, recording the operations performed on all the tensors. - The resulting recording of a standalone function produces - ScriptFunction. + ``ScriptFunction``. -- The resulting recording of nn.Module.forward or nn.Module produces - ScriptModule. +- The resulting recording of ``nn.Module.forward`` or ``nn.Module`` + produces ``ScriptModule``. In the same way like scripted model, traced model can be used for inference or saved on disk using ``torch.jit.save`` function and after @@ -667,8 +718,9 @@ original PyTorch model code definitions. 5: hamper - 2.35% -Benchmark Traced Model Inference -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Benchmark Traced Model Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -679,14 +731,12 @@ Benchmark Traced Model Inference .. parsed-literal:: - 12.7 ms ± 39.6 µs per loop (mean ± std. dev. of 7 runs, 10 loops each) + 12.7 ms ± 61.1 µs per loop (mean ± std. dev. of 7 runs, 10 loops each) Convert PyTorch Traced Model to OpenVINO Intermediate Representation -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ - -The conversion step for a traced model to OpenVINO IR is similar to the -original PyTorch model. +`⇑ <#top>`__ The conversion step for a traced model to OpenVINO IR is +similar to the original PyTorch model. .. code:: ipython3 @@ -723,7 +773,7 @@ original PyTorch model. Benchmark OpenVINO Model Inference Converted From Traced Model -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +`⇑ <#top>`__ .. code:: ipython3 @@ -734,5 +784,5 @@ Benchmark OpenVINO Model Inference Converted From Traced Model .. parsed-literal:: - 3.08 ms ± 9.07 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) + 3.05 ms ± 6.85 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) diff --git a/docs/notebooks/102-pytorch-to-openvino-with-output_files/index.html b/docs/notebooks/102-pytorch-to-openvino-with-output_files/index.html index 20bdcce272a..cbb9df8a8b2 100644 --- a/docs/notebooks/102-pytorch-to-openvino-with-output_files/index.html +++ b/docs/notebooks/102-pytorch-to-openvino-with-output_files/index.html @@ -1,20 +1,20 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/102-pytorch-to-openvino-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/102-pytorch-to-openvino-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/102-pytorch-to-openvino-with-output_files/


../
-102-pytorch-to-openvino-with-output_11_0.jpg       12-Jul-2023 00:11               54874
-102-pytorch-to-openvino-with-output_11_0.png       12-Jul-2023 00:11              542516
-102-pytorch-to-openvino-with-output_20_0.jpg       12-Jul-2023 00:11               54874
-102-pytorch-to-openvino-with-output_20_0.png       12-Jul-2023 00:11              542516
-102-pytorch-to-openvino-with-output_31_0.jpg       12-Jul-2023 00:11               54874
-102-pytorch-to-openvino-with-output_31_0.png       12-Jul-2023 00:11              542516
-102-pytorch-to-openvino-with-output_35_0.jpg       12-Jul-2023 00:11               54874
-102-pytorch-to-openvino-with-output_35_0.png       12-Jul-2023 00:11              542516
-102-pytorch-to-openvino-with-output_39_0.jpg       12-Jul-2023 00:11               54874
-102-pytorch-to-openvino-with-output_39_0.png       12-Jul-2023 00:11              542516
-102-pytorch-to-openvino-with-output_43_0.jpg       12-Jul-2023 00:11               54874
-102-pytorch-to-openvino-with-output_43_0.png       12-Jul-2023 00:11              542516
-102-pytorch-to-openvino-with-output_47_0.jpg       12-Jul-2023 00:11               54874
-102-pytorch-to-openvino-with-output_47_0.png       12-Jul-2023 00:11              542516
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/102-pytorch-to-openvino-with-output_files/


../
+102-pytorch-to-openvino-with-output_11_0.jpg       16-Aug-2023 01:31               54874
+102-pytorch-to-openvino-with-output_11_0.png       16-Aug-2023 01:31              542516
+102-pytorch-to-openvino-with-output_20_0.jpg       16-Aug-2023 01:31               54874
+102-pytorch-to-openvino-with-output_20_0.png       16-Aug-2023 01:31              542516
+102-pytorch-to-openvino-with-output_31_0.jpg       16-Aug-2023 01:31               54874
+102-pytorch-to-openvino-with-output_31_0.png       16-Aug-2023 01:31              542516
+102-pytorch-to-openvino-with-output_35_0.jpg       16-Aug-2023 01:31               54874
+102-pytorch-to-openvino-with-output_35_0.png       16-Aug-2023 01:31              542516
+102-pytorch-to-openvino-with-output_39_0.jpg       16-Aug-2023 01:31               54874
+102-pytorch-to-openvino-with-output_39_0.png       16-Aug-2023 01:31              542516
+102-pytorch-to-openvino-with-output_43_0.jpg       16-Aug-2023 01:31               54874
+102-pytorch-to-openvino-with-output_43_0.png       16-Aug-2023 01:31              542516
+102-pytorch-to-openvino-with-output_47_0.jpg       16-Aug-2023 01:31               54874
+102-pytorch-to-openvino-with-output_47_0.png       16-Aug-2023 01:31              542516
 

diff --git a/docs/notebooks/103-paddle-to-openvino-classification-with-output.rst b/docs/notebooks/103-paddle-to-openvino-classification-with-output.rst index 19423995938..72b335538b1 100644 --- a/docs/notebooks/103-paddle-to-openvino-classification-with-output.rst +++ b/docs/notebooks/103-paddle-to-openvino-classification-with-output.rst @@ -1,12 +1,14 @@ Convert a PaddlePaddle Model to OpenVINO™ IR ============================================ +.. _top: + This notebook shows how to convert a MobileNetV3 model from `PaddleHub `__, pre-trained on the `ImageNet `__ dataset, to OpenVINO IR. It also shows how to perform classification inference on a sample image, using `OpenVINO -Runtime `__ +Runtime `__ and compares the results of the `PaddlePaddle `__ model with the IR model. @@ -14,11 +16,28 @@ IR model. Source of the `model `__. -Preparation ------------ +**Table of contents**: + +- `Preparation <#1preparation>`__ + + - `Imports <#imports>`__ + - `Settings <#settings>`__ + +- `Show Inference on PaddlePaddle Model <#show-inference-on-paddlepaddle-model>`__ +- `Convert the Model to OpenVINO IR Format <#convert-the-model-to-openvino-ir-format>`__ +- `Select inference device <#select-inference-device>`__ +- `Show Inference on OpenVINO Model <#show-inference-on-openvino-model>`__ +- `Timing and Comparison <#timing-and-comparison>`__ +- `Select inference device <#select-inference-device>`__ +- `References <#references>`__ + +Preparation `⇑ <#top>`__ +############################################################################################################################### + + +Imports `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Imports -~~~~~~~ .. code:: ipython3 @@ -61,18 +80,19 @@ Imports .. parsed-literal:: - 2023-07-11 22:26:05 INFO: Loading faiss with AVX2 support. - 2023-07-11 22:26:05 INFO: Successfully loaded faiss with AVX2 support. + 2023-08-15 22:28:07 INFO: Loading faiss with AVX2 support. + 2023-08-15 22:28:07 INFO: Successfully loaded faiss with AVX2 support. -Settings -~~~~~~~~ +Settings `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Set ``IMAGE_FILENAME`` to the filename of an image to use. Set ``MODEL_NAME`` to the PaddlePaddle model to download from PaddleHub. ``MODEL_NAME`` will also be the base name for the IR model. The notebook is tested with the -`mobilenet_v3_large_x1_0 `__ +`MobileNetV3_large_x1_0 `__ model. Other models may use different preprocessing methods and therefore require some modification to get the same results on the original and converted model. @@ -110,8 +130,9 @@ PaddleHub. This may take a while. Model Extracted to "./model". -Show Inference on PaddlePaddle Model ------------------------------------- +Show Inference on PaddlePaddle Model `⇑ <#top>`__ +############################################################################################################################### + In the next cell, we load the model, load and display an image, do inference on that image, and then show the top three prediction results. @@ -130,7 +151,7 @@ inference on that image, and then show the top three prediction results. .. parsed-literal:: - [2023/07/11 22:26:25] ppcls WARNING: The current running environment does not support the use of GPU. CPU has been used instead. + [2023/08/15 22:28:34] ppcls WARNING: The current running environment does not support the use of GPU. CPU has been used instead. Labrador retriever, 0.75138 German short-haired pointer, 0.02373 Great Dane, 0.01848 @@ -182,8 +203,8 @@ the same method. It is useful to show the output of the ``process_image()`` function, to see the effect of cropping and resizing. Because of the normalization, -the colors will look strange, and matplotlib will warn about clipping -values. +the colors will look strange, and ``matplotlib`` will warn about +clipping values. .. code:: ipython3 @@ -196,7 +217,7 @@ values. .. parsed-literal:: - 2023-07-11 22:26:25 WARNING: Clipping input data to the valid range for imshow with RGB data ([0..1] for floats or [0..255] for integers). + 2023-08-15 22:28:34 WARNING: Clipping input data to the valid range for imshow with RGB data ([0..1] for floats or [0..255] for integers). .. parsed-literal:: @@ -208,7 +229,7 @@ values. .. parsed-literal:: - + @@ -232,8 +253,9 @@ OpenVINO model. partition = line.split("\n")[0].partition(" ") class_id_map[int(partition[0])] = str(partition[-1]) -Convert the Model to OpenVINO IR Format ---------------------------------------- +Convert the Model to OpenVINO IR Format `⇑ <#top>`__ +############################################################################################################################### + Call the OpenVINO Model Optimizer Python API to convert the PaddlePaddle model to OpenVINO IR, with FP32 precision. ``mo.convert_model`` function @@ -256,10 +278,11 @@ for more information about Model Optimizer. else: print(f"{model_xml} already exists.") -Select inference device ------------------------ +Select inference device `⇑ <#top>`__ +############################################################################################################################### -select device from dropdown list for running inference using OpenVINO + +Select device from dropdown list for running inference using OpenVINO: .. code:: ipython3 @@ -284,8 +307,9 @@ select device from dropdown list for running inference using OpenVINO -Show Inference on OpenVINO Model --------------------------------- +Show Inference on OpenVINO Model `⇑ <#top>`__ +############################################################################################################################### + Load the IR model, get model information, load the image, do inference, convert the inference to a meaningful result, and show the output. See @@ -332,8 +356,9 @@ information. .. image:: 103-paddle-to-openvino-classification-with-output_files/103-paddle-to-openvino-classification-with-output_23_1.png -Timing and Comparison ---------------------- +Timing and Comparison `⇑ <#top>`__ +############################################################################################################################### + Measure the time it takes to do inference on fifty images and compare the result. The timing information gives an indication of performance. @@ -386,7 +411,7 @@ Note that many optimizations are possible to improve the performance. .. parsed-literal:: - PaddlePaddle model on CPU: 0.0074 seconds per image, FPS: 135.48 + PaddlePaddle model on CPU: 0.0071 seconds per image, FPS: 141.47 PaddlePaddle result: Labrador retriever, 0.75138 @@ -400,10 +425,11 @@ Note that many optimizations are possible to improve the performance. .. image:: 103-paddle-to-openvino-classification-with-output_files/103-paddle-to-openvino-classification-with-output_27_1.png -Select inference device ------------------------ +Select inference device `⇑ <#top>`__ +############################################################################################################################### -select device from dropdown list for running inference using OpenVINO + +Select device from dropdown list for running inference using OpenVINO: .. code:: ipython3 @@ -448,7 +474,7 @@ select device from dropdown list for running inference using OpenVINO .. parsed-literal:: - OpenVINO IR model in OpenVINO Runtime (AUTO): 0.0031 seconds per image, FPS: 326.50 + OpenVINO IR model in OpenVINO Runtime (AUTO): 0.0030 seconds per image, FPS: 337.97 OpenVINO result: Labrador retriever, 0.75138 @@ -462,8 +488,9 @@ select device from dropdown list for running inference using OpenVINO .. image:: 103-paddle-to-openvino-classification-with-output_files/103-paddle-to-openvino-classification-with-output_30_1.png -References ----------- +References `⇑ <#top>`__ +############################################################################################################################### + - `PaddleClas `__ - `OpenVINO PaddlePaddle diff --git a/docs/notebooks/103-paddle-to-openvino-classification-with-output_files/index.html b/docs/notebooks/103-paddle-to-openvino-classification-with-output_files/index.html index 788123f477f..ca87d2ee17e 100644 --- a/docs/notebooks/103-paddle-to-openvino-classification-with-output_files/index.html +++ b/docs/notebooks/103-paddle-to-openvino-classification-with-output_files/index.html @@ -1,11 +1,11 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/103-paddle-to-openvino-classification-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/103-paddle-to-openvino-classification-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/103-paddle-to-openvino-classification-with-output_files/


../
-103-paddle-to-openvino-classification-with-outp..> 12-Jul-2023 00:11              120883
-103-paddle-to-openvino-classification-with-outp..> 12-Jul-2023 00:11              224886
-103-paddle-to-openvino-classification-with-outp..> 12-Jul-2023 00:11              224886
-103-paddle-to-openvino-classification-with-outp..> 12-Jul-2023 00:11              224886
-103-paddle-to-openvino-classification-with-outp..> 12-Jul-2023 00:11              224886
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/103-paddle-to-openvino-classification-with-output_files/


../
+103-paddle-to-openvino-classification-with-outp..> 16-Aug-2023 01:31              120883
+103-paddle-to-openvino-classification-with-outp..> 16-Aug-2023 01:31              224886
+103-paddle-to-openvino-classification-with-outp..> 16-Aug-2023 01:31              224886
+103-paddle-to-openvino-classification-with-outp..> 16-Aug-2023 01:31              224886
+103-paddle-to-openvino-classification-with-outp..> 16-Aug-2023 01:31              224886
 

diff --git a/docs/notebooks/104-model-tools-with-output.rst b/docs/notebooks/104-model-tools-with-output.rst index b00d78901ca..c60c541cece 100644 --- a/docs/notebooks/104-model-tools-with-output.rst +++ b/docs/notebooks/104-model-tools-with-output.rst @@ -1,37 +1,58 @@ Working with Open Model Zoo Models ================================== +.. _top: + This tutorial shows how to download a model from `Open Model Zoo `__, convert it to OpenVINO™ IR format, show information about the model, and benchmark -the model. +the model. + +**Table of contents**: + +- `OpenVINO and Open Model Zoo Tools <#openvino-and-open-model-zoo-tools>`__ +- `Preparation <#preparation>`__ + + - `Model Name <#model-name>`__ + - `Imports <#imports>`__ + - `Settings and Configuration <#settings-and-configuration>`__ + +- `Download a Model from Open Model Zoo <#download-a-model-from-open-model-zoo>`__ +- `Convert a Model to OpenVINO IR format <#convert-a-model-to-openvino-ir-format>`__ +- `Get Model Information <#get-model-information>`__ +- `Run Benchmark Tool <#run-benchmark-tool>`__ + + - `Benchmark with Different Settings <#benchmark-with-different-settings>`__ + +OpenVINO and Open Model Zoo Tools `⇑ <#top>`__ +############################################################################################################################### -OpenVINO and Open Model Zoo Tools ---------------------------------- OpenVINO and Open Model Zoo tools are listed in the table below. +------------+--------------+-----------------------------------------+ | Tool | Command | Description | +============+==============+=========================================+ -| Model | omz_download | Download models from Open Model Zoo. | -| Downloader | er | | +| Model | ``omz_downlo | Download models from Open Model Zoo. | +| Downloader | ader`` | | +------------+--------------+-----------------------------------------+ -| Model | omz_converte | Convert Open Model Zoo models to | -| Converter | r | OpenVINO’s IR format. | +| Model | ``omz_conver | Convert Open Model Zoo models to | +| Converter | ter`` | OpenVINO’s IR format. | +------------+--------------+-----------------------------------------+ -| Info | omz_info_dum | Print information about Open Model Zoo | -| Dumper | per | models. | +| Info | ``omz_info_d | Print information about Open Model Zoo | +| Dumper | umper`` | models. | +------------+--------------+-----------------------------------------+ -| Benchmark | benchmark_ap | Benchmark model performance by | -| Tool | p | computing inference time. | +| Benchmark | ``benchmark_ | Benchmark model performance by | +| Tool | app`` | computing inference time. | +------------+--------------+-----------------------------------------+ -Preparation ------------ +Preparation `⇑ <#top>`__ +############################################################################################################################### + + +Model Name `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Model Name -~~~~~~~~~~ Set ``model_name`` to the name of the Open Model Zoo model to use in this notebook. Refer to the list of @@ -46,8 +67,9 @@ pre-trained models for a full list of models that can be used. Set # model_name = "resnet-50-pytorch" model_name = "mobilenet-v2-pytorch" -Imports -~~~~~~~ +Imports `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -61,8 +83,9 @@ Imports sys.path.append("../utils") from notebook_utils import DeviceNotFoundAlert, NotebookAlert -Settings and Configuration -~~~~~~~~~~~~~~~~~~~~~~~~~~ +Settings and Configuration `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Set the file and directory paths. By default, this notebook downloads models from Open Model Zoo to the ``open_model_zoo_models`` directory in @@ -100,8 +123,9 @@ The following settings can be changed: base_model_dir: model, omz_cache_dir: cache, gpu_availble: False -Download a Model from Open Model Zoo ------------------------------------- +Download a Model from Open Model Zoo `⇑ <#top>`__ +############################################################################################################################### + Specify, display and run the Model Downloader command to download the model. @@ -140,8 +164,9 @@ Downloading mobilenet-v2-pytorch… -Convert a Model to OpenVINO IR format -------------------------------------- +Convert a Model to OpenVINO IR format `⇑ <#top>`__ +############################################################################################################################### + Specify, display and run the Model Converter command to convert the model to OpenVINO IR format. Model conversion may take a while. The @@ -177,25 +202,26 @@ Converting mobilenet-v2-pytorch… .. parsed-literal:: ========== Converting mobilenet-v2-pytorch to ONNX - Conversion to ONNX command: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/bin/python -- /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/model_zoo/internal_scripts/pytorch_to_onnx.py --model-name=mobilenet_v2 --weights=model/public/mobilenet-v2-pytorch/mobilenet_v2-b0353104.pth --import-module=torchvision.models --input-shape=1,3,224,224 --output-file=model/public/mobilenet-v2-pytorch/mobilenet-v2.onnx --input-names=data --output-names=prob + Conversion to ONNX command: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/bin/python -- /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/model_zoo/internal_scripts/pytorch_to_onnx.py --model-name=mobilenet_v2 --weights=model/public/mobilenet-v2-pytorch/mobilenet_v2-b0353104.pth --import-module=torchvision.models --input-shape=1,3,224,224 --output-file=model/public/mobilenet-v2-pytorch/mobilenet-v2.onnx --input-names=data --output-names=prob ONNX check passed successfully. ========== Converting mobilenet-v2-pytorch to IR (FP16) - Conversion command: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/bin/python -- /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/bin/mo --framework=onnx --output_dir=/tmp/tmpe5yh3lmf --model_name=mobilenet-v2-pytorch --input=data '--mean_values=data[123.675,116.28,103.53]' '--scale_values=data[58.624,57.12,57.375]' --reverse_input_channels --output=prob --input_model=model/public/mobilenet-v2-pytorch/mobilenet-v2.onnx '--layout=data(NCHW)' '--input_shape=[1, 3, 224, 224]' --compress_to_fp16=True + Conversion command: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/bin/python -- /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/bin/mo --framework=onnx --output_dir=/tmp/tmp3q4nxrwu --model_name=mobilenet-v2-pytorch --input=data '--mean_values=data[123.675,116.28,103.53]' '--scale_values=data[58.624,57.12,57.375]' --reverse_input_channels --output=prob --input_model=model/public/mobilenet-v2-pytorch/mobilenet-v2.onnx '--layout=data(NCHW)' '--input_shape=[1, 3, 224, 224]' --compress_to_fp16=True [ INFO ] Generated IR will be compressed to FP16. If you get lower accuracy, please consider disabling compression by removing argument --compress_to_fp16 or set it to false --compress_to_fp16=False. Find more information about compression to FP16 at https://docs.openvino.ai/latest/openvino_docs_MO_DG_FP16_Compression.html [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/latest/openvino_2_0_transition_guide.html [ SUCCESS ] Generated IR version 11 model. - [ SUCCESS ] XML file: /tmp/tmpe5yh3lmf/mobilenet-v2-pytorch.xml - [ SUCCESS ] BIN file: /tmp/tmpe5yh3lmf/mobilenet-v2-pytorch.bin + [ SUCCESS ] XML file: /tmp/tmp3q4nxrwu/mobilenet-v2-pytorch.xml + [ SUCCESS ] BIN file: /tmp/tmp3q4nxrwu/mobilenet-v2-pytorch.bin -Get Model Information ---------------------- +Get Model Information `⇑ <#top>`__ +############################################################################################################################### + The Info Dumper prints the following information for Open Model Zoo models: @@ -240,8 +266,8 @@ information in a dictionary. 'description': 'MobileNet V2 is image classification model pre-trained on ImageNet dataset. This is a PyTorch* implementation of MobileNetV2 architecture as described in the paper "Inverted Residuals and Linear Bottlenecks: Mobile Networks for Classification, Detection and Segmentation" .\nThe model input is a blob that consists of a single image of "1, 3, 224, 224" in "RGB" order.\nThe model output is typical object classifier for the 1000 different classifications matching with those in the ImageNet database.', 'framework': 'pytorch', 'license_url': 'https://raw.githubusercontent.com/pytorch/vision/master/LICENSE', - 'accuracy_config': '/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/model_zoo/models/public/mobilenet-v2-pytorch/accuracy-check.yml', - 'model_config': '/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/model_zoo/models/public/mobilenet-v2-pytorch/model.yml', + 'accuracy_config': '/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/model_zoo/models/public/mobilenet-v2-pytorch/accuracy-check.yml', + 'model_config': '/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/model_zoo/models/public/mobilenet-v2-pytorch/model.yml', 'precisions': ['FP16', 'FP32'], 'quantization_output_precisions': ['FP16-INT8', 'FP32-INT8'], 'subdirectory': 'public/mobilenet-v2-pytorch', @@ -273,8 +299,9 @@ file. model/public/mobilenet-v2-pytorch/FP16/mobilenet-v2-pytorch.xml exists: True -Run Benchmark Tool ------------------- +Run Benchmark Tool `⇑ <#top>`__ +############################################################################################################################### + By default, Benchmark Tool runs inference for 60 seconds in asynchronous mode on CPU. It returns inference speed as latency (milliseconds per @@ -321,7 +348,7 @@ seconds… [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.THROUGHPUT. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 30.53 ms + [ INFO ] Read model took 29.61 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] data (node: data) : f32 / [N,C,H,W] / [1,3,224,224] @@ -335,7 +362,7 @@ seconds… [ INFO ] Model outputs: [ INFO ] prob (node: prob) : f32 / [...] / [1,1000] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 149.89 ms + [ INFO ] Compile model took 154.76 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: [ INFO ] NETWORK_NAME: torch_jit @@ -357,21 +384,22 @@ seconds… [ INFO ] Fill input 'data' with random values [Step 10/11] Measuring performance (Start inference asynchronously, 6 inference requests, limits: 15000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 6.66 ms + [ INFO ] First inference took 7.60 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 20136 iterations - [ INFO ] Duration: 15007.99 ms + [ INFO ] Count: 20076 iterations + [ INFO ] Duration: 15004.20 ms [ INFO ] Latency: - [ INFO ] Median: 4.33 ms - [ INFO ] Average: 4.34 ms - [ INFO ] Min: 3.20 ms - [ INFO ] Max: 11.84 ms - [ INFO ] Throughput: 1341.69 FPS + [ INFO ] Median: 4.34 ms + [ INFO ] Average: 4.35 ms + [ INFO ] Min: 2.53 ms + [ INFO ] Max: 11.71 ms + [ INFO ] Throughput: 1338.03 FPS -Benchmark with Different Settings -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Benchmark with Different Settings `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The ``benchmark_app`` tool displays logging information that is not always necessary. A more compact result is achieved when the output is @@ -456,9 +484,9 @@ Benchmark command: command ended Traceback (most recent call last): - File "/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/tools/benchmark/main.py", line 327, in main + File "/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/tools/benchmark/main.py", line 327, in main benchmark.set_allow_auto_batching(False) - File "/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/tools/benchmark/benchmark.py", line 63, in set_allow_auto_batching + File "/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/tools/benchmark/benchmark.py", line 63, in set_allow_auto_batching self.core.set_property({'ALLOW_AUTO_BATCHING': flag}) RuntimeError: Check 'false' failed at src/inference/src/core.cpp:238: @@ -484,9 +512,9 @@ Benchmark command: Check 'false' failed at src/plugins/auto/src/plugin_config.cpp:55: property: ALLOW_AUTO_BATCHING: not supported Traceback (most recent call last): - File "/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/tools/benchmark/main.py", line 327, in main + File "/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/tools/benchmark/main.py", line 327, in main benchmark.set_allow_auto_batching(False) - File "/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/tools/benchmark/benchmark.py", line 63, in set_allow_auto_batching + File "/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/tools/benchmark/benchmark.py", line 63, in set_allow_auto_batching self.core.set_property({'ALLOW_AUTO_BATCHING': flag}) RuntimeError: Check 'false' failed at src/inference/src/core.cpp:238: Check 'false' failed at src/plugins/auto/src/plugin_config.cpp:55: diff --git a/docs/notebooks/105-language-quantize-bert-with-output.rst b/docs/notebooks/105-language-quantize-bert-with-output.rst index 7875c6b4c2d..f3d9df156fd 100644 --- a/docs/notebooks/105-language-quantize-bert-with-output.rst +++ b/docs/notebooks/105-language-quantize-bert-with-output.rst @@ -1,6 +1,8 @@ Quantize NLP models with Post-Training Quantization ​in NNCF ============================================================ +.. _top: + This tutorial demonstrates how to apply ``INT8`` quantization to the Natural Language Processing model known as `BERT `__, using @@ -22,12 +24,27 @@ and datasets. It consists of the following steps: - Compare the performance of the original, converted and quantized models. +**Table of contents**: + +- `Imports <#imports>`__ +- `Settings <#settings>`__ +- `Prepare the Model <#prepare-the-model>`__ +- `Prepare the Dataset <#prepare-the-dataset>`__ +- `Optimize model using NNCF Post-training Quantization API <#optimize-model-using-nncf-post-training-quantization-api>`__ +- `Load and Test OpenVINO Model <#load-and-test-openvino-model>`__ + + - `Select inference device <#select-inference-device>`__ + +- `Compare F1-score of FP32 and INT8 models <#compare-f1-score-of-fp32-and-int8-models>`__ +- `Compare Performance of the Original, Converted and Quantized Models <#compare-performance-of-the-original,-converted-and-quantized-models>`__ + .. code:: ipython3 !pip install -q "nncf>=2.5.0" datasets evaluate -Imports -------- +Imports `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -56,10 +73,10 @@ Imports .. parsed-literal:: - 2023-07-11 22:27:10.887837: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. - 2023-07-11 22:27:10.921844: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + 2023-08-15 22:29:19.942802: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-08-15 22:29:19.975605: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. - 2023-07-11 22:27:11.494944: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + 2023-08-15 22:29:20.517786: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT .. parsed-literal:: @@ -67,8 +84,9 @@ Imports INFO:nncf:NNCF initialized successfully. Supported frameworks detected: torch, tensorflow, onnx, openvino -Settings --------- +Settings `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -82,13 +100,15 @@ Settings os.makedirs(DATA_DIR, exist_ok=True) os.makedirs(MODEL_DIR, exist_ok=True) -Prepare the Model ------------------ +Prepare the Model `⇑ <#top>`__ +############################################################################################################################### + Perform the following: -- Download and unpack pre-trained BERT model for MRPC by PyTorch. -- Convert the model to the OpenVINO Intermediate Representation (OpenVINO IR) +- Download and unpack pre-trained BERT model for MRPC by PyTorch. +- Convert the model to the OpenVINO Intermediate Representation + (OpenVINO IR) .. code:: ipython3 @@ -107,12 +127,12 @@ Convert the original PyTorch model to the OpenVINO Intermediate Representation. From OpenVINO 2023.0, we can directly convert a model from the PyTorch -format to the OpenVINO IR format using Model Optimizer. Following +format to the OpenVINO IR format using model conversion API. Following PyTorch model formats are supported: -- torch.nn.Module -- torch.jit.ScriptModule -- torch.jit.ScriptFunction +- ``torch.nn.Module`` +- ``torch.jit.ScriptModule`` +- ``torch.jit.ScriptFunction`` .. code:: ipython3 @@ -142,17 +162,16 @@ PyTorch model formats are supported: .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/jit/annotations.py:309: UserWarning: TorchScript will treat type annotations of Tensor dtype-specific subtypes as if they are normal Tensors. dtype constraints are not enforced in compilation either. + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/jit/annotations.py:309: UserWarning: TorchScript will treat type annotations of Tensor dtype-specific subtypes as if they are normal Tensors. dtype constraints are not enforced in compilation either. warnings.warn("TorchScript will treat type annotations of Tensor " -Prepare the Dataset -------------------- +Prepare the Dataset `⇑ <#top>`__ +############################################################################################################################### -We download the `General Language Understanding Evaluation -(GLUE) `__ dataset for the MRPC task from -HuggingFace datasets. Then, we tokenize the data with a pre-trained BERT -tokenizer from HuggingFace. +We download the `General Language Understanding Evaluation (GLUE) `__ dataset +for the MRPC task from HuggingFace datasets. Then, we tokenize the data +with a pre-trained BERT tokenizer from HuggingFace. .. code:: ipython3 @@ -171,15 +190,9 @@ tokenizer from HuggingFace. data_source = create_data_source() +Optimize model using NNCF Post-training Quantization API `⇑ <#top>`__ +############################################################################################################################### -.. parsed-literal:: - - Found cached dataset glue (/opt/home/k8sworker/.cache/huggingface/datasets/glue/mrpc/1.0.0/dacbe3125aa31d7f70367a07a8a9e72a5a0bfeb5fc42e75c9db75b96da6053ad) - Loading cached processed dataset at /opt/home/k8sworker/.cache/huggingface/datasets/glue/mrpc/1.0.0/dacbe3125aa31d7f70367a07a8a9e72a5a0bfeb5fc42e75c9db75b96da6053ad/cache-b5f4c739eb2a4a9f.arrow - - -Optimize model using NNCF Post-training Quantization API --------------------------------------------------------- `NNCF `__ provides a suite of advanced algorithms for Neural Networks inference optimization in @@ -387,8 +400,8 @@ The optimization process contains the following steps: .. parsed-literal:: - Statistics collection: 100%|██████████| 300/300 [00:24<00:00, 12.02it/s] - Biases correction: 100%|██████████| 74/74 [00:25<00:00, 2.94it/s] + Statistics collection: 100%|██████████| 300/300 [00:24<00:00, 12.04it/s] + Biases correction: 100%|██████████| 74/74 [00:25<00:00, 2.95it/s] .. code:: ipython3 @@ -396,20 +409,22 @@ The optimization process contains the following steps: compressed_model_xml = Path(MODEL_DIR) / "quantized_bert_mrpc.xml" ov.serialize(quantized_model, compressed_model_xml) -Load and Test OpenVINO Model ----------------------------- +Load and Test OpenVINO Model `⇑ <#top>`__ +############################################################################################################################### + To load and test converted model, perform the following: -* Load the model and compile it for selected device. -* Prepare the input. -* Run the inference. -* Get the answer from the model output. +- Load the model and compile it for selected device. +- Prepare the input. +- Run the inference. +- Get the answer from the model output. -Select inference device -~~~~~~~~~~~~~~~~~~~~~~~ +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -select device from dropdown list for running inference using OpenVINO + +Select device from dropdown list for running inference using OpenVINO: .. code:: ipython3 @@ -465,8 +480,9 @@ changing ``sample_idx`` to another value (from 0 to 407). The same meaning: yes -Compare F1-score of FP32 and INT8 models ----------------------------------------- +Compare F1-score of FP32 and INT8 models `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -509,8 +525,8 @@ Compare F1-score of FP32 and INT8 models F1 score: 0.8995 -Compare Performance of the Original, Converted and Quantized Models -------------------------------------------------------------------- +Compare Performance of the Original, Converted and Quantized Models. `⇑ <#top>`__ +############################################################################################################################### Compare the original PyTorch model with OpenVINO converted and quantized models (``FP32``, ``INT8``) to see the difference in performance. It is @@ -562,14 +578,19 @@ Frames Per Second (FPS) for images. .. parsed-literal:: - PyTorch model on CPU: 0.071 seconds per sentence, SPS: 14.09 - IR FP32 model in OpenVINO Runtime/AUTO: 0.022 seconds per sentence, SPS: 45.98 - OpenVINO IR INT8 model in OpenVINO Runtime/AUTO: 0.010 seconds per sentence, SPS: 98.77 + We strongly recommend passing in an `attention_mask` since your input_ids may be padded. See https://huggingface.co/docs/transformers/troubleshooting#incorrect-output-when-padding-tokens-arent-masked. + + +.. parsed-literal:: + + PyTorch model on CPU: 0.070 seconds per sentence, SPS: 14.22 + IR FP32 model in OpenVINO Runtime/AUTO: 0.021 seconds per sentence, SPS: 48.42 + OpenVINO IR INT8 model in OpenVINO Runtime/AUTO: 0.010 seconds per sentence, SPS: 98.01 Finally, measure the inference performance of OpenVINO ``FP32`` and -``INT8`` models. For this purpose, use `Benchmark -Tool `__ +``INT8`` models. For this purpose, use +`Benchmark Tool `__ in OpenVINO. **Note**: The ``benchmark_app`` tool is able to measure the @@ -600,9 +621,9 @@ in OpenVINO. [ ERROR ] Check 'false' failed at src/inference/src/core.cpp:84: Device with "device" name is not registered in the OpenVINO Runtime Traceback (most recent call last): - File "/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/tools/benchmark/main.py", line 103, in main + File "/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/tools/benchmark/main.py", line 103, in main benchmark.print_version_info() - File "/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/tools/benchmark/benchmark.py", line 48, in print_version_info + File "/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/tools/benchmark/benchmark.py", line 48, in print_version_info for device, version in self.core.get_versions(self.device).items(): RuntimeError: Check 'false' failed at src/inference/src/core.cpp:84: Device with "device" name is not registered in the OpenVINO Runtime @@ -628,9 +649,9 @@ in OpenVINO. [ ERROR ] Check 'false' failed at src/inference/src/core.cpp:84: Device with "device" name is not registered in the OpenVINO Runtime Traceback (most recent call last): - File "/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/tools/benchmark/main.py", line 103, in main + File "/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/tools/benchmark/main.py", line 103, in main benchmark.print_version_info() - File "/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/tools/benchmark/benchmark.py", line 48, in print_version_info + File "/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/tools/benchmark/benchmark.py", line 48, in print_version_info for device, version in self.core.get_versions(self.device).items(): RuntimeError: Check 'false' failed at src/inference/src/core.cpp:84: Device with "device" name is not registered in the OpenVINO Runtime diff --git a/docs/notebooks/106-auto-device-with-output.rst b/docs/notebooks/106-auto-device-with-output.rst index 5ef4b0a22ce..be32615035c 100644 --- a/docs/notebooks/106-auto-device-with-output.rst +++ b/docs/notebooks/106-auto-device-with-output.rst @@ -1,6 +1,8 @@ Automatic Device Selection with OpenVINO™ ========================================= +.. _top: + The `Auto device `__ (or AUTO in short) selects the most suitable device for inference by @@ -25,10 +27,36 @@ immediately on the CPU and then transparently shifts inference to the GPU, once it is ready. This dramatically reduces the time to execute first inference. -.. image:: https://camo.githubusercontent.com/cc526c3f5fc992cc7176d097894303248adbd04b4d158bd98e65edc8270af5fc/68747470733a2f2f757365722d696d616765732e67697468756275736572636f6e74656e742e636f6d2f31353730393732332f3136313435313834372d37353965326264622d373062632d343633642d393831382d3430306330636366336331362e706e67 +.. figure:: https://user-images.githubusercontent.com/15709723/161451847-759e2bdb-70bc-463d-9818-400c0ccf3c16.png + :alt: auto + + auto + +**Table of contents**: + +- `Import modules and create Core <#import-modules-and-create-core>`__ +- `Convert the model to OpenVINO IR format <#convert-the-model-to-openvino-ir-format>`__ +- `(1) Simplify selection logic <#1-simplify-selection-logic>`__ + + - `Default behavior of Core::compile_model API without device_name <#default-behavior-of-core::compile_model-api-without-device_name>`__ + - `Explicitly pass AUTO as device_name to Core::compile_model API <#explicitly-pass-auto-as-device_name-to-core::compile_model-api>`__ + +- `(2) Improve the first inference latency <#2-improve-the-first-inference-latency>`__ + + - `Load an Image <#load-an-image>`__ + - `Load the model to GPU device and perform inference <#load-the-model-to-gpu-device-and-perform-inference>`__ + - `Load the model using AUTO device and do inference <#load-the-model-using-auto-device-and-do-inference>`__ + +- `(3) Achieve different performance for different targets <#3-achieve-different-performance-for-different-targets>`__ + + - `Class and callback definition <#class-and-callback-definition>`__ + - `Inference with THROUGHPUT hint <#inference-with-throughput-hint>`__ + - `Inference with LATENCY hint <#inference-with-latency-hint>`__ + - `Difference in FPS and latency <#difference-in-fps-and-latency>`__ + +Import modules and create Core `⇑ <#top>`__ +############################################################################################################################### -Import modules and create Core ------------------------------- .. code:: ipython3 @@ -50,28 +78,27 @@ Import modules and create Core device to have meaningful results. -Convert the model to OpenVINO IR format ---------------------------------------- +Convert the model to OpenVINO IR format `⇑ <#top>`__ +############################################################################################################################### + This tutorial uses `resnet50 `__ model from `torchvision `__ library. ResNet 50 is image classification model pre-trained on ImageNet -dataset described in paper `“Deep Residual Learning for Image -Recognition” `__. From OpenVINO +dataset described in paper `“Deep Residual Learning for Image Recognition” `__. From OpenVINO 2023.0, we can directly convert a model from the PyTorch format to the -OpenVINO IR format using Model Optimizer. To convert model, we should -provide model object instance into ``mo.convert_model`` function, +OpenVINO IR format using model conversion API. To convert model, we +should provide model object instance into ``mo.convert_model`` function, optionally, we can specify input shape for conversion (by default models from PyTorch converted with dynamic input shapes). ``mo.convert_model`` -returns openvino.runtime.Model object ready to be loaded on device with -``openvino.runtime.Core().compile_model`` or serialized for next usage -with ``openvino.runtime.serialize``. +returns openvino.runtime.Model object ready to be loaded on a device +with ``openvino.runtime.Core().compile_model`` or serialized for next +usage with ``openvino.runtime.serialize``. -For more information about Model Optimizer, see the `Model Optimizer -Developer -Guide `__. +For more information about model conversion API, see this +`page `__. .. code:: ipython3 @@ -99,14 +126,15 @@ Guide `__ +############################################################################################################################### + -Default behavior of Core::compile_model API without device_name -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Default behavior of Core::compile_model API without device_name `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -By default, ``compile_model`` API will select **AUTO** as -``device_name`` if no device is specified. +By default, ``compile_model`` API will select **AUTO** as ``device_name`` if no +device is specified. .. code:: ipython3 @@ -137,11 +165,11 @@ By default, ``compile_model`` API will select **AUTO** as Deleted compiled_model -Explicitly pass AUTO as device_name to Core::compile_model API -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Explicitly pass AUTO as device_name to Core::compile_model API `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -It is optional, but passing AUTO explicitly as ``device_name`` may -improve readability of your code. +It is optional, but passing AUTO explicitly as +``device_name`` may improve readability of your code. .. code:: ipython3 @@ -171,13 +199,13 @@ improve readability of your code. Deleted compiled_model -(2) Improve the first inference latency ---------------------------------------- +(2) Improve the first inference latency `⇑ <#top>`__ +############################################################################################################################### -One of the benefits of using AUTO device selection is reducing FIL -(first inference latency). FIL is the model compilation time combined -with the first inference execution time. Using the CPU device explicitly -will produce the shortest first inference latency, as the OpenVINO graph +One of the benefits of using AUTO device selection is reducing FIL (first inference +latency). FIL is the model compilation time combined with the first +inference execution time. Using the CPU device explicitly will produce +the shortest first inference latency, as the OpenVINO graph representation loads quickly on CPU, using just-in-time (JIT) compilation. The challenge is with GPU devices since OpenCL graph complication to GPU-optimized kernels takes a few seconds to complete. @@ -185,11 +213,12 @@ This initialization time may be intolerable for some applications. To avoid this delay, the AUTO uses CPU transparently as the first inference device until GPU is ready. -Load an Image -~~~~~~~~~~~~~ +Load an Image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -torchvision library provides model specific input transformation -function, we will reuse it for preparing input data. +Torchvision library provides model specific +input transformation function, we will reuse it for preparing input +data. .. code:: ipython3 @@ -209,8 +238,9 @@ function, we will reuse it for preparing input data. -Load the model to GPU device and perform inference -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Load the model to GPU device and perform inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -236,8 +266,8 @@ Load the model to GPU device and perform inference A GPU device is not available. Available devices are: ['CPU'] -Load the model using AUTO device and do inference -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Load the model using AUTO device and do inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ When GPU is the best available device, the first few inferences will be executed on CPU until GPU is ready. @@ -268,8 +298,8 @@ executed on CPU until GPU is ready. # Deleted model will wait for compiling on the selected device to complete. del compiled_model -(3) Achieve different performance for different targets -------------------------------------------------------- +(3) Achieve different performance for different targets `⇑ <#top>`__ +############################################################################################################################### It is an advantage to define **performance hints** when using Automatic Device Selection. By specifying a **THROUGHPUT** or **LATENCY** hint, @@ -280,14 +310,13 @@ hints do not require any device-specific settings and they are completely portable between devices – meaning AUTO can configure the performance hint on whichever device is being used. -For more information, refer to the `Performance -Hints `__ -section of `Automatic Device -Selection `__ +For more information, refer to the `Performance Hints `__ +section of `Automatic Device Selection `__ article. -Class and callback definition -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Class and callback definition `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -385,8 +414,9 @@ Class and callback definition metrics_update_interval = 10 metrics_update_num = 6 -Inference with THROUGHPUT hint -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Inference with THROUGHPUT hint `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Loop for inference and update the FPS/Latency every @metrics_update_interval seconds. @@ -424,17 +454,18 @@ Loop for inference and update the FPS/Latency every Compiling Model for AUTO device with THROUGHPUT hint Start inference, 6 groups of FPS/latency will be measured over 10s intervals - throughput: 190.70fps, latency: 29.76ms, time interval: 10.02s - throughput: 191.95fps, latency: 30.48ms, time interval: 10.00s - throughput: 192.78fps, latency: 30.40ms, time interval: 10.00s - throughput: 191.39fps, latency: 30.62ms, time interval: 10.00s - throughput: 192.18fps, latency: 30.44ms, time interval: 10.03s - throughput: 191.33fps, latency: 30.62ms, time interval: 10.00s + throughput: 189.24fps, latency: 30.04ms, time interval: 10.00s + throughput: 192.12fps, latency: 30.48ms, time interval: 10.01s + throughput: 191.27fps, latency: 30.64ms, time interval: 10.00s + throughput: 190.87fps, latency: 30.69ms, time interval: 10.01s + throughput: 189.50fps, latency: 30.89ms, time interval: 10.02s + throughput: 190.30fps, latency: 30.79ms, time interval: 10.01s Done -Inference with LATENCY hint -~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Inference with LATENCY hint `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Loop for inference and update the FPS/Latency for each @metrics_update_interval seconds @@ -473,17 +504,18 @@ Loop for inference and update the FPS/Latency for each Compiling Model for AUTO Device with LATENCY hint Start inference, 6 groups fps/latency will be out with 10s interval - throughput: 136.99fps, latency: 6.75ms, time interval: 10.00s - throughput: 140.91fps, latency: 6.74ms, time interval: 10.01s - throughput: 140.83fps, latency: 6.74ms, time interval: 10.00s - throughput: 140.90fps, latency: 6.74ms, time interval: 10.00s - throughput: 140.83fps, latency: 6.74ms, time interval: 10.00s - throughput: 140.85fps, latency: 6.74ms, time interval: 10.00s + throughput: 138.76fps, latency: 6.68ms, time interval: 10.00s + throughput: 141.79fps, latency: 6.70ms, time interval: 10.00s + throughput: 142.39fps, latency: 6.68ms, time interval: 10.00s + throughput: 142.30fps, latency: 6.68ms, time interval: 10.00s + throughput: 142.30fps, latency: 6.68ms, time interval: 10.01s + throughput: 142.53fps, latency: 6.67ms, time interval: 10.00s Done -Difference in FPS and latency -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Difference in FPS and latency `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 diff --git a/docs/notebooks/106-auto-device-with-output_files/106-auto-device-with-output_25_0.png b/docs/notebooks/106-auto-device-with-output_files/106-auto-device-with-output_25_0.png index d3aecde3701..ec960eef224 100644 --- a/docs/notebooks/106-auto-device-with-output_files/106-auto-device-with-output_25_0.png +++ b/docs/notebooks/106-auto-device-with-output_files/106-auto-device-with-output_25_0.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:03224259b8d1a7dd982c9d5b2baae8bdca495cbe008f83bc5b0a0fd3ecc47ab6 -size 27172 +oid sha256:7858463aadc74dfe6af6906da15e208305234a13a8463c4f4e4632630fffde70 +size 27107 diff --git a/docs/notebooks/106-auto-device-with-output_files/106-auto-device-with-output_26_0.png b/docs/notebooks/106-auto-device-with-output_files/106-auto-device-with-output_26_0.png index 3e6b0edb53f..415a474e73b 100644 --- a/docs/notebooks/106-auto-device-with-output_files/106-auto-device-with-output_26_0.png +++ b/docs/notebooks/106-auto-device-with-output_files/106-auto-device-with-output_26_0.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:d962a84b77491e4c1f9bdcb4e0c58f6aae55f11f8e59ae747077b166e706fd76 -size 40994 +oid sha256:48a1f70dac1b326af8f205247fc377c9fb72286526c2d19507ada4b520a779ed +size 39987 diff --git a/docs/notebooks/106-auto-device-with-output_files/index.html b/docs/notebooks/106-auto-device-with-output_files/index.html index 27cbe80da0b..a32b6ac7960 100644 --- a/docs/notebooks/106-auto-device-with-output_files/index.html +++ b/docs/notebooks/106-auto-device-with-output_files/index.html @@ -1,10 +1,10 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/106-auto-device-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/106-auto-device-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/106-auto-device-with-output_files/


../
-106-auto-device-with-output_12_0.jpg               12-Jul-2023 00:11              121563
-106-auto-device-with-output_12_0.png               12-Jul-2023 00:11              869661
-106-auto-device-with-output_25_0.png               12-Jul-2023 00:11               27172
-106-auto-device-with-output_26_0.png               12-Jul-2023 00:11               40994
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/106-auto-device-with-output_files/


../
+106-auto-device-with-output_12_0.jpg               16-Aug-2023 01:31              121563
+106-auto-device-with-output_12_0.png               16-Aug-2023 01:31              869661
+106-auto-device-with-output_25_0.png               16-Aug-2023 01:31               27107
+106-auto-device-with-output_26_0.png               16-Aug-2023 01:31               39987
 

diff --git a/docs/notebooks/107-speech-recognition-quantization-data2vec-with-output.rst b/docs/notebooks/107-speech-recognition-quantization-data2vec-with-output.rst index fce865b55f7..c79f51edaab 100644 --- a/docs/notebooks/107-speech-recognition-quantization-data2vec-with-output.rst +++ b/docs/notebooks/107-speech-recognition-quantization-data2vec-with-output.rst @@ -1,6 +1,8 @@ Quantize Speech Recognition Models using NNCF PTQ API ===================================================== +.. _top: + This tutorial demonstrates how to use the NNCF (Neural Network Compression Framework) 8-bit quantization in post-training mode (without the fine-tuning pipeline) to optimize the speech recognition model, @@ -19,8 +21,24 @@ steps: - Compare performance of the original and quantized models. - Compare Accuracy of the Original and Quantized Models. -Download and prepare model --------------------------- +**Table of contents**: + +- `Download and prepare model <#download-and-prepare-model>`__ + + - `Obtain Pytorch model representation <#obtain-pytorch-model-representation>`__ + - `Convert model to OpenVINO Intermediate Representation <#convert-model-to-openvino-intermediate-representation>`__ + - `Prepare inference data <#prepare-inference-data>`__ + +- `Check model inference result <#check-model-inference-result>`__ +- `Validate model accuracy on dataset <#validate-model-accuracy-on-dataset>`__ +- `Quantization <#quantization>`__ +- `Check INT8 model inference result <#check-int8-model-inference-result>`__ +- `Compare Performance of the Original and Quantized Models <#compare-performance-of-the-original-and-quantized-models>`__ +- `Compare Accuracy of the Original and Quantized Models <#compare-accuracy-of-the-original-and-quantized-models>`__ + +Download and prepare model `⇑ <#top>`__ +############################################################################################################################### + data2vec is a framework for self-supervised representation learning for images, speech, and text as described in `data2vec: A General Framework @@ -38,8 +56,9 @@ In our case, we will use ``data2vec-audio-base-960h`` model, which was finetuned on 960 hours of audio from LibriSpeech Automatic Speech Recognition corpus and distributed as part of HuggingFace transformers. -Obtain Pytorch model representation -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Obtain Pytorch model representation `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + For instantiating PyTorch model class, we should use ``Data2VecAudioForCTC.from_pretrained`` method with providing model ID @@ -53,7 +72,7 @@ model specific pre- and post-processing steps. .. code:: ipython3 - !pip install -q 'openvino-dev>=2023.0.0' 'nncf>=2.5.0' + !pip install -q "openvino-dev>=2023.0.0" "nncf>=2.5.0" !pip install -q soundfile librosa transformers onnx .. code:: ipython3 @@ -63,8 +82,9 @@ model specific pre- and post-processing steps. processor = Wav2Vec2Processor.from_pretrained("facebook/data2vec-audio-base-960h") model = Data2VecAudioForCTC.from_pretrained("facebook/data2vec-audio-base-960h") -Convert model to OpenVINO Intermediate Representation -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Convert model to OpenVINO Intermediate Representation `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -134,11 +154,12 @@ Convert model to OpenVINO Intermediate Representation Read IR model from model/data2vec-audo-base.xml -Prepare inference data -~~~~~~~~~~~~~~~~~~~~~~ +Prepare inference data `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + For demonstration purposes, we will use short dummy version of -librispeach dataset - ``patrickvonplaten/librispeech_asr_dummy`` to +LibriSpeech dataset - ``patrickvonplaten/librispeech_asr_dummy`` to speed up model evaluation. Model accuracy can be different from reported in the paper. For reproducing original accuracy, use ``librispeech_asr`` dataset. @@ -174,8 +195,9 @@ dataset. Loading cached processed dataset at /home/adrian/.cache/huggingface/datasets/patrickvonplaten___librispeech_asr_dummy/clean/2.1.0/f2c70a4d03ab4410954901bde48c54b85ca1b7f9bf7d616e7e2a72b5ee6ddbfc/cache-4e0f4916cd205b24.arrow -Check model inference result ----------------------------- +Check model inference result `⇑ <#top>`__ +############################################################################################################################### + The code below is used for running model inference on a single sample from the dataset. It contains the following steps: @@ -247,8 +269,9 @@ For reference, see the same function provided for OpenVINO model. -Validate model accuracy on dataset ----------------------------------- +Validate model accuracy on dataset `⇑ <#top>`__ +############################################################################################################################### + For model accuracy evaluation, `Word Error Rate `__ metric can be @@ -307,8 +330,9 @@ library. [OpenVino] Word Error Rate: 0.0383 -Quantization ------------- +Quantization `⇑ <#top>`__ +############################################################################################################################### + `NNCF `__ provides a suite of advanced algorithms for Neural Networks inference optimization in @@ -318,9 +342,11 @@ Create a quantized model from the pre-trained ``FP16`` model and the calibration dataset. The optimization process contains the following steps: -1. Create a Dataset for quantization. -2. Run ``nncf.quantize`` for getting an optimized model. The ``nncf.quantize`` function provides an interface for model quantization. It requires an instance of the OpenVINO Model and quantization dataset. Optionally, some additional parameters for the configuration quantization process (number of samples for quantization, preset, ignored scope, etc.) can be provided. For more accurate results, we should keep the operation in the postprocessing subgraph in floating point precision, using the ``ignored_scope`` parameter. ``advanced_parameters`` can be used to specify advanced quantization parameters for fine-tuning the quantization algorithm. In this tutorial we pass range estimator parameters for activations. For more information, see `Tune quantization parameters `__. -3. Serialize OpenVINO IR model using ``openvino.runtime.serialize`` function. +:: + + 1. Create a Dataset for quantization. + 2. Run `nncf.quantize` for getting an optimized model. The `nncf.quantize` function provides an interface for model quantization. It requires an instance of the OpenVINO Model and quantization dataset. Optionally, some additional parameters for the configuration quantization process (number of samples for quantization, preset, ignored scope, etc.) can be provided. For more accurate results, we should keep the operation in the postprocessing subgraph in floating point precision, using the `ignored_scope` parameter. `advanced_parameters` can be used to specify advanced quantization parameters for fine-tuning the quantization algorithm. In this tutorial we pass range estimator parameters for activations. For more information see [Tune quantization parameters](https://docs.openvino.ai/2023.0/basic_quantization_flow.html#tune-quantization-parameters). + 3. Serialize OpenVINO IR model using `openvino.runtime.serialize` function. .. code:: ipython3 @@ -589,8 +615,9 @@ saved using ``serialize`` function. quantized_model_path = Path(f"{MODEL_NAME}_openvino_model/{MODEL_NAME}_quantized.xml") serialize(quantized_model, str(quantized_model_path)) -Check INT8 model inference result ---------------------------------- +Check INT8 model inference result `⇑ <#top>`__ +############################################################################################################################### + ``INT8`` model is the same in usage like the original one. We need to read it, using the ``core.read_model`` method and load on the device, @@ -628,8 +655,8 @@ using ``core.compile_model``. After that, we can reuse the same -Compare Performance of the Original and Quantized Models --------------------------------------------------------- +Compare Performance of the Original and Quantized Models `⇑ <#top>`__ +############################################################################################################################### `Benchmark Tool `__ @@ -791,8 +818,9 @@ is used to measure the inference performance of the ``FP16`` and [ INFO ] Throughput: 38.24 FPS -Compare Accuracy of the Original and Quantized Models ------------------------------------------------------ +Compare Accuracy of the Original and Quantized Models `⇑ <#top>`__ +############################################################################################################################### + Finally, calculate WER metric for the ``INT8`` model representation and compare it with the ``FP16`` result. diff --git a/docs/notebooks/108-gpu-device-with-output.rst b/docs/notebooks/108-gpu-device-with-output.rst new file mode 100644 index 00000000000..27617889028 --- /dev/null +++ b/docs/notebooks/108-gpu-device-with-output.rst @@ -0,0 +1,1375 @@ +Working with GPUs in OpenVINO™ +============================== + +.. _top: + +**Table of contents**: + +- `Introduction <#introduction>`__ + + - `Install required packages <#install-required-packages>`__ + +- `Checking GPUs with Query Device <#checking-gpus-with-query-device>`__ + + - `List GPUs with core.available_devices <#list-gpus-with-core.available_devices>`__ + - `Check Properties with core.get_property <#check-properties-with-core.get_property>`__ + - `Brief Descriptions of Key Properties <#brief-descriptions-of-key-properties>`__ + +- `Compiling a Model on GPU <#compiling-a-model-on-gpu>`__ + + - `Download and Convert a Model <#download-and-convert-a-model>`__ + + - `Download and unpack the Model <#download-and-unpack-the-model>`__ + - `Convert the Model to OpenVINO IR format <#convert-the-model-to-openvino-ir-format>`__ + + - `Compile with Default Configuration <#compile-with-default-configuration>`__ + - `Reduce Compile Time through Model Caching <#reduce-compile-time-through-model-caching>`__ + - `Throughput and Latency Performance Hints <#throughput-and-latency-performance-hints>`__ + - `Using Multiple GPUs with Multi-Device and Cumulative Throughput <#using-multiple-gpus-with-multi-device-and-cumulative-throughput>`__ + +- `Performance Comparison with benchmark_app <#performance-comparison-with-benchmark_app>`__ +- `CPU vs GPU with Latency Hint <#cpu-vs-gpu-with-latency-hint>`__ +- `CPU vs GPU with Throughput Hint <#cpu-vs-gpu-with-throughput-hint>`__ +- `Single GPU vs Multiple GPUs <#single-gpu-vs-multiple-gpus>`__ +- `Basic Application Using GPUs <#basic-application-using-gpus>`__ + + - `Import Necessary Packages <#import-necessary-packages>`__ + - `Compile the Model <#compile-the-model>`__ + - `Load and Preprocess Video Frames <#load-and-preprocess-video-frames>`__ + - `Define Model Output Classes <#define-model-output-classes>`__ + - `Set up Asynchronous Pipeline <#set-up-asynchronous-pipeline>`__ + + - `Callback Definition <#callback-definition>`__ + - `Create Async Pipeline <#create-async-pipeline>`__ + + - `Perform Inference <#perform-inference>`__ + - `Process Results <#process-results>`__ +- `Conclusion <#conclusion>`__ + +This tutorial provides a high-level overview of working with Intel GPUs +in OpenVINO. It shows how to use Query Device to list system GPUs and +check their properties, and it explains some of the key properties. It +shows how to compile a model on GPU with performance hints and how to +use multiple GPUs using MULTI or CUMULATIVE_THROUGHPUT. + +The tutorial also shows example commands for benchmark_app that can be +run to compare GPU performance in different configurations. It also +provides the code for a basic end-to-end application that compiles a +model on GPU and uses it to run inference. + +Introduction `⇑ <#top>`__ +############################################################################################################################### + + +Originally, graphic processing units (GPUs) began as specialized chips, +developed to accelerate the rendering of computer graphics. In contrast +to CPUs, which have few but powerful cores, GPUs have many more +specialized cores, making them ideal for workloads that can be +parallelized into simpler tasks. Nowadays, one such workload is deep +learning, where GPUs can easily accelerate inference of neural networks +by splitting operations across multiple cores. + +OpenVINO supports inference on Intel integrated GPUs (which are included +with most `Intel® Core™ desktop and mobile +processors `__) +or on Intel discrete GPU products like the `Intel® Arc™ A-Series +Graphics +cards `__ +and `Intel® Data Center GPU Flex +Series `__. +To get started, first `install +OpenVINO `__ +on a system equipped with one or more Intel GPUs. Follow the `GPU +configuration +instructions `__ +to configure OpenVINO to work with your GPU. Then, read on to learn how +to accelerate inference with GPUs in OpenVINO! + +Install required packages `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +.. code:: ipython3 + + !pip install -q "openvino-dev>=2023.0.0" + !pip install -q tensorflow + + # Fetch `notebook_utils` module + import urllib.request + urllib.request.urlretrieve( + url='https://raw.githubusercontent.com/openvinotoolkit/openvino_notebooks/main/notebooks/utils/notebook_utils.py', + filename='notebook_utils.py' + ) + + + + +.. parsed-literal:: + + ('notebook_utils.py', ) + + + +Checking GPUs with Query Device `⇑ <#top>`__ +############################################################################################################################### + + +In this section, we will see how to list the available GPUs and check +their properties. Some of the key properties will also be defined. + +List GPUs with core.available_devices `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +OpenVINO Runtime provides the ``available_devices`` method for checking +which devices are available for inference. The following code will +output a list of compatible OpenVINO devices, in which Intel GPUs should +appear. + +.. code:: ipython3 + + from openvino.runtime import Core + + core = Core() + core.available_devices + + + + +.. parsed-literal:: + + ['CPU', 'GPU'] + + + +Note that GPU devices are numbered starting at 0, where the integrated +GPU always takes the id ``0`` if the system has one. For instance, if +the system has a CPU, an integrated and discrete GPU, we should expect +to see a list like this: ``['CPU', 'GPU.0', 'GPU.1']``. To simplify its +use, the “GPU.0” can also be addressed with just “GPU”. For more +details, see the `Device Naming +Convention `__ +section. + +If the GPUs are installed correctly on the system and still do not +appear in the list, follow the steps described +`here `__ +to configure your GPU drivers to work with OpenVINO. Once we have the +GPUs working with OpenVINO, we can proceed with the next sections. + +Check Properties with core.get_property `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +To get information about the GPUs, we can use device properties. In +OpenVINO, devices have properties that describe their characteristics +and configuration. Each property has a name and associated value that +can be queried with the ``get_property`` method. + +To get the value of a property, such as the device name, we can use the +``get_property`` method as follows: + +.. code:: ipython3 + + device = "GPU" + + core.get_property(device, "FULL_DEVICE_NAME") + + + + +.. parsed-literal:: + + 'Intel(R) Graphics [0x46a6] (iGPU)' + + + +Each device also has a specific property called +``SUPPORTED_PROPERTIES``, that enables viewing all the available +properties in the device. We can check the value for each property by +simply looping through the dictionary returned by +``core.get_property("GPU", "SUPPORTED_PROPERTIES")`` and then querying +for that property. + +.. code:: ipython3 + + print(f"{device} SUPPORTED_PROPERTIES:\n") + supported_properties = core.get_property(device, "SUPPORTED_PROPERTIES") + indent = len(max(supported_properties, key=len)) + + for property_key in supported_properties: + if property_key not in ('SUPPORTED_METRICS', 'SUPPORTED_CONFIG_KEYS', 'SUPPORTED_PROPERTIES'): + try: + property_val = core.get_property(device, property_key) + except TypeError: + property_val = 'UNSUPPORTED TYPE' + print(f"{property_key:<{indent}}: {property_val}") + + +.. parsed-literal:: + + GPU SUPPORTED_PROPERTIES: + + AVAILABLE_DEVICES : ['0'] + RANGE_FOR_ASYNC_INFER_REQUESTS: (1, 2, 1) + RANGE_FOR_STREAMS : (1, 2) + OPTIMAL_BATCH_SIZE : 1 + MAX_BATCH_SIZE : 1 + CACHING_PROPERTIES : {'GPU_UARCH_VERSION': 'RO', 'GPU_EXECUTION_UNITS_COUNT': 'RO', 'GPU_DRIVER_VERSION': 'RO', 'GPU_DEVICE_ID': 'RO'} + DEVICE_ARCHITECTURE : GPU: v12.0.0 + FULL_DEVICE_NAME : Intel(R) Graphics [0x46a6] (iGPU) + DEVICE_UUID : UNSUPPORTED TYPE + DEVICE_TYPE : Type.INTEGRATED + DEVICE_GOPS : UNSUPPORTED TYPE + OPTIMIZATION_CAPABILITIES : ['FP32', 'BIN', 'FP16', 'INT8'] + GPU_DEVICE_TOTAL_MEM_SIZE : UNSUPPORTED TYPE + GPU_UARCH_VERSION : 12.0.0 + GPU_EXECUTION_UNITS_COUNT : 96 + GPU_MEMORY_STATISTICS : UNSUPPORTED TYPE + PERF_COUNT : False + MODEL_PRIORITY : Priority.MEDIUM + GPU_HOST_TASK_PRIORITY : Priority.MEDIUM + GPU_QUEUE_PRIORITY : Priority.MEDIUM + GPU_QUEUE_THROTTLE : Priority.MEDIUM + GPU_ENABLE_LOOP_UNROLLING : True + CACHE_DIR : + PERFORMANCE_HINT : PerformanceMode.UNDEFINED + COMPILATION_NUM_THREADS : 20 + NUM_STREAMS : 1 + PERFORMANCE_HINT_NUM_REQUESTS : 0 + INFERENCE_PRECISION_HINT : + DEVICE_ID : 0 + + +Brief Descriptions of Key Properties `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Each device has several properties as seen in the last command. Some of +the key properties are: + +- ``FULL_DEVICE_NAME`` - The product name of the GPU and whether it is + an integrated or discrete GPU (iGPU or dGPU). +- ``OPTIMIZATION_CAPABILITIES`` - The model data types (INT8, FP16, + FP32, etc) that are supported by this GPU. +- ``GPU_EXECUTION_UNITS_COUNT`` - The execution cores available in the + GPU’s architecture, which is a relative measure of the GPU’s + processing power. +- ``RANGE_FOR_STREAMS`` - The number of processing streams available on + the GPU that can be used to execute parallel inference requests. When + compiling a model in LATENCY or THROUGHPUT mode, OpenVINO will + automatically select the best number of streams for low latency or + high throughput. +- ``PERFORMANCE_HINT`` - A high-level way to tune the device for a + specific performance metric, such as latency or throughput, without + worrying about device-specific settings. +- ``CACHE_DIR`` - The directory where the model cache data is stored to + speed up compilation time. + +To learn more about devices and properties, see the `Query Device +Properties `__ +page. + +Compiling a Model on GPU `⇑ <#top>`__ +############################################################################################################################### + + +Now, we know how to list the GPUs in the system and check their +properties. We can easily use one for compiling and running models with +OpenVINO `GPU +plugin `__. + +Download and Convert a Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +This tutorial uses the ``ssdlite_mobilenet_v2`` model. The +``ssdlite_mobilenet_v2`` model is used for object detection. The model +was trained on `Common Objects in Context +(COCO) `__ dataset version with 91 +categories of object. For details, see the +`paper `__. + +Download and unpack the Model `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + + +Use the ``download_file`` function from the ``notebook_utils`` to +download an archive with the model. It automatically creates a directory +structure and downloads the selected model. This step is skipped if the +package is already downloaded. + +.. code:: ipython3 + + import sys + import tarfile + from pathlib import Path + + sys.path.append("../utils") + + import notebook_utils as utils + + # A directory where the model will be downloaded. + base_model_dir = Path("./model").expanduser() + + model_name = "ssdlite_mobilenet_v2" + archive_name = Path(f"{model_name}_coco_2018_05_09.tar.gz") + + # Download the archive + downloaded_model_path = base_model_dir / archive_name + if not downloaded_model_path.exists(): + model_url = f"http://download.tensorflow.org/models/object_detection/{archive_name}" + utils.download_file(model_url, downloaded_model_path.name, downloaded_model_path.parent) + + # Unpack the model + tf_model_path = base_model_dir / archive_name.with_suffix("").stem / "frozen_inference_graph.pb" + if not tf_model_path.exists(): + with tarfile.open(downloaded_model_path) as file: + file.extractall(base_model_dir) + + + +.. parsed-literal:: + + model/ssdlite_mobilenet_v2_coco_2018_05_09.tar.gz: 0%| | 0.00/48.7M [00:00`__ +------------------------------------------------------------------------------------------------------------------------------- + + +To convert the model to OpenVINO IR with ``FP16`` precision, use model +conversion API. The models are saved to the ``model/ir_model/`` +directory. For more details about model conversion, see this +`page `__. + +.. code:: ipython3 + + from openvino.tools import mo + from openvino.runtime import serialize + from openvino.tools.mo.front import tf as ov_tf_front + + precision = 'FP16' + + # The output path for the conversion. + model_path = base_model_dir / 'ir_model' / f'{model_name}_{precision.lower()}.xml' + + trans_config_path = Path(ov_tf_front.__file__).parent / "ssd_v2_support.json" + pipeline_config = base_model_dir / archive_name.with_suffix("").stem / "pipeline.config" + + model = None + if not model_path.exists(): + model = mo.convert_model(input_model=tf_model_path, + input_shape=[1, 300, 300, 3], + layout='NHWC', + compress_to_fp16=True if precision == 'FP16' else False, + transformations_config=trans_config_path, + tensorflow_object_detection_api_pipeline_config=pipeline_config, + reverse_input_channels=True) + serialize(model, str(model_path)) + print("IR model saved to {}".format(model_path)) + else: + print("Read IR model from {}".format(model_path)) + model = core.read_model(model_path) + + +.. parsed-literal:: + + [ WARNING ] The Preprocessor block has been removed. Only nodes performing mean value subtraction and scaling (if applicable) are kept. + + +.. parsed-literal:: + + IR model saved to model/ir_model/ssdlite_mobilenet_v2_fp16.xml + + +Compile with Default Configuration `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +When the model is ready, first we need to read it, using the +``read_model`` method. Then, we can use the ``compile_model`` method and +specify the name of the device we want to compile the model on, in this +case, “GPU”. + +.. code:: ipython3 + + compiled_model = core.compile_model(model, device) + +If you have multiple GPUs in the system, you can specify which one to +use by using “GPU.0”, “GPU.1”, etc. Any of the device names returned by +the ``available_devices`` method are valid device specifiers. You may +also use “AUTO”, which will automatically select the best device for +inference (which is often the GPU). To learn more about AUTO plugin, +visit the `Automatic Device +Selection `__ +page as well as the `AUTO device +tutorial `__. + +Reduce Compile Time through Model Caching `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Depending on the model used, device-specific optimizations and network +compilations can cause the compile step to be time-consuming, especially +with larger models, which may lead to bad user experience in the +application, in which they are used. To solve this, OpenVINO can cache +the model once it is compiled on supported devices and reuse it in later +``compile_model`` calls by simply setting a cache folder beforehand. For +instance, to cache the same model we compiled above, we can do the +following: + +.. code:: ipython3 + + import time + from pathlib import Path + + # Create cache folder + cache_folder = Path("cache") + cache_folder.mkdir(exist_ok=True) + + start = time.time() + core = Core() + + # Set cache folder + core.set_property({'CACHE_DIR': cache_folder}) + + # Compile the model as before + model = core.read_model(model=model_path) + compiled_model = core.compile_model(model, device) + print(f"Cache enabled (first time) - compile time: {time.time() - start}s") + + +.. parsed-literal:: + + Cache enabled (first time) - compile time: 1.692436695098877s + + +To get an idea of the effect that caching can have, we can measure the +compile times with caching enabled and disabled as follows: + +.. code:: ipython3 + + start = time.time() + core = Core() + core.set_property({'CACHE_DIR': 'cache'}) + model = core.read_model(model=model_path) + compiled_model = core.compile_model(model, device) + print(f"Cache enabled - compile time: {time.time() - start}s") + + start = time.time() + core = Core() + model = core.read_model(model=model_path) + compiled_model = core.compile_model(model, device) + print(f"Cache disabled - compile time: {time.time() - start}s") + + +.. parsed-literal:: + + Cache enabled - compile time: 0.26888394355773926s + Cache disabled - compile time: 1.982884168624878s + + +The actual time improvements will depend on the environment as well as +the model being used but it is definitely something to consider when +optimizing an application. To read more about this, see the `Model +Caching `__ +docs. + +Throughput and Latency Performance Hints `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +To simplify device and pipeline configuration, OpenVINO provides +high-level performance hints that automatically set the batch size and +number of parallel threads to use for inference. The “LATENCY” +performance hint optimizes for fast inference times while the +“THROUGHPUT” performance hint optimizes for high overall bandwidth or +FPS. + +To use the “LATENCY” performance hint, add +``{"PERFORMANCE_HINT": "LATENCY"}`` when compiling the model as shown +below. For GPUs, this automatically minimizes the batch size and number +of parallel streams such that all of the compute resources can focus on +completing a single inference as fast as possible. + +.. code:: ipython3 + + compiled_model = core.compile_model(model, device, {"PERFORMANCE_HINT": "LATENCY"}) + +To use the “THROUGHPUT” performance hint, add +``{"PERFORMANCE_HINT": "THROUGHPUT"}`` when compiling the model. For +GPUs, this creates multiple processing streams to efficiently utilize +all the execution cores and optimizes the batch size to fill the +available memory. + +.. code:: ipython3 + + compiled_model = core.compile_model(model, device, {"PERFORMANCE_HINT": "THROUGHPUT"}) + +Using Multiple GPUs with Multi-Device and Cumulative Throughput `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + +The latency and throughput hints mentioned above are great and can make +a difference when used adequately but they usually use just one device, +either due to the `AUTO +plugin `__ +or by manual specification of the device name as above. When we have +multiple devices, such as an integrated and discrete GPU, we may use +both at the same time to improve the utilization of the resources. In +order to do this, OpenVINO provides a virtual device called +`MULTI `__, +which is just a combination of the existent devices that knows how to +split inference work between them, leveraging the capabilities of each +device. + +As an example, if we want to use both integrated and discrete GPUs and +the CPU at the same time, we can compile the model as follows: + +``compiled_model = core.compile_model(model=model, device_name="MULTI:GPU.1,GPU.0,CPU")`` + +Note that we always need to explicitly specify the device list for MULTI +to work, otherwise MULTI does not know which devices are available for +inference. However, this is not the only way to use multiple devices in +OpenVINO. There is another performance hint called +“CUMULATIVE_THROUGHPUT” that works similar to MULTI, except it uses the +devices automatically selected by AUTO. This way, we do not need to +manually specify devices to use. Below is an example showing how to use +“CUMULATIVE_THROUGHPUT”, equivalent to the MULTI one: + +``compiled_model = core.compile_model(model=model, device_name="AUTO", config={"PERFORMANCE_HINT": "CUMULATIVE_THROUGHPUT"})`` + + **Important**: **The “THROUGHPUT”, “MULTI”, and + “CUMULATIVE_THROUGHPUT” modes are only applicable to asynchronous + inferencing pipelines. The example at the end of this article shows + how to set up an asynchronous pipeline that takes advantage of + parallelism to increase throughput.** To learn more, see + `Asynchronous + Inferencing `__ + in OpenVINO as well as the `Asynchronous Inference + notebook `__. + +Performance Comparison with benchmark_app `⇑ <#top>`__ +############################################################################################################################### + + +Given all the different options available when compiling a model, it may +be difficult to know which settings work best for a certain application. +Thankfully, OpenVINO provides ``benchmark_app`` - a performance +benchmarking tool. + +The basic syntax of ``benchmark_app`` is as follows: + +``benchmark_app -m PATH_TO_MODEL -d TARGET_DEVICE -hint {throughput,cumulative_throughput,latency,none}`` + +where ``TARGET_DEVICE`` is any device shown by the ``available_devices`` +method as well as the MULTI and AUTO devices we saw previously, and the +value of hint should be one of the values between brackets. + +Note that benchmark_app only requires the model path to run but both the +device and hint arguments will be useful to us. For more advanced +usages, the tool itself has other options that can be checked by running +``benchmark_app -h`` or reading the +`docs `__. +The following example shows how to benchmark a simple model, using a GPU +with a latency focus: + +.. code:: ipython3 + + !benchmark_app -m {model_path} -d GPU -hint latency + + +.. parsed-literal:: + + [Step 1/11] Parsing and validating input arguments + [ INFO ] Parsing input parameters + [Step 2/11] Loading OpenVINO Runtime + [ INFO ] OpenVINO: + [ INFO ] Build ................................. 2022.3.0-9052-9752fafe8eb-releases/2022/3 + [ INFO ] + [ INFO ] Device info: + [ INFO ] GPU + [ INFO ] Build ................................. 2022.3.0-9052-9752fafe8eb-releases/2022/3 + [ INFO ] + [ INFO ] + [Step 3/11] Setting device configuration + [Step 4/11] Reading model files + [ INFO ] Loading model files + [ INFO ] Read model took 14.02 ms + [ INFO ] Original model I/O parameters: + [ INFO ] Model inputs: + [ INFO ] image_tensor , image_tensor:0 (node: image_tensor) : u8 / [N,H,W,C] / [1,300,300,3] + [ INFO ] Model outputs: + [ INFO ] detection_boxes:0 (node: DetectionOutput) : f32 / [...] / [1,1,100,7] + [Step 5/11] Resizing model to match image sizes and given batch + [ INFO ] Model batch size: 1 + [Step 6/11] Configuring input of the model + [ INFO ] Model inputs: + [ INFO ] image_tensor , image_tensor:0 (node: image_tensor) : u8 / [N,H,W,C] / [1,300,300,3] + [ INFO ] Model outputs: + [ INFO ] detection_boxes:0 (node: DetectionOutput) : f32 / [...] / [1,1,100,7] + [Step 7/11] Loading the model to the device + [ INFO ] Compile model took 1932.50 ms + [Step 8/11] Querying optimal runtime parameters + [ INFO ] Model: + [ INFO ] NETWORK_NAME: frozen_inference_graph + [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 1 + [ INFO ] PERF_COUNT: False + [ INFO ] MODEL_PRIORITY: Priority.MEDIUM + [ INFO ] GPU_HOST_TASK_PRIORITY: Priority.MEDIUM + [ INFO ] GPU_QUEUE_PRIORITY: Priority.MEDIUM + [ INFO ] GPU_QUEUE_THROTTLE: Priority.MEDIUM + [ INFO ] GPU_ENABLE_LOOP_UNROLLING: True + [ INFO ] CACHE_DIR: + [ INFO ] PERFORMANCE_HINT: PerformanceMode.LATENCY + [ INFO ] COMPILATION_NUM_THREADS: 20 + [ INFO ] NUM_STREAMS: 1 + [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 + [ INFO ] INFERENCE_PRECISION_HINT: + [ INFO ] DEVICE_ID: 0 + [Step 9/11] Creating infer requests and preparing input tensors + [ WARNING ] No input files were given for input 'image_tensor'!. This input will be filled with random values! + [ INFO ] Fill input 'image_tensor' with random values + [Step 10/11] Measuring performance (Start inference asynchronously, 1 inference requests, limits: 60000 ms duration) + [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). + [ INFO ] First inference took 6.17 ms + [Step 11/11] Dumping statistics report + [ INFO ] Count: 12710 iterations + [ INFO ] Duration: 60006.58 ms + [ INFO ] Latency: + [ INFO ] Median: 4.52 ms + [ INFO ] Average: 4.57 ms + [ INFO ] Min: 3.13 ms + [ INFO ] Max: 17.62 ms + [ INFO ] Throughput: 211.81 FPS + + +For completeness, let us list here some of the comparisons we may want +to do by varying the device and hint used. Note that the actual +performance may depend on the hardware used. Generally, we should expect +GPU to be better than CPU, whereas multiple GPUs should be better than a +single GPU as long as there is enough work for each of them. + +CPU vs GPU with Latency Hint `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + + +.. code:: ipython3 + + !benchmark_app -m {model_path} -d CPU -hint latency + + +.. parsed-literal:: + + [Step 1/11] Parsing and validating input arguments + [ INFO ] Parsing input parameters + [Step 2/11] Loading OpenVINO Runtime + [ INFO ] OpenVINO: + [ INFO ] Build ................................. 2022.3.0-9052-9752fafe8eb-releases/2022/3 + [ INFO ] + [ INFO ] Device info: + [ INFO ] CPU + [ INFO ] Build ................................. 2022.3.0-9052-9752fafe8eb-releases/2022/3 + [ INFO ] + [ INFO ] + [Step 3/11] Setting device configuration + [Step 4/11] Reading model files + [ INFO ] Loading model files + [ INFO ] Read model took 30.38 ms + [ INFO ] Original model I/O parameters: + [ INFO ] Model inputs: + [ INFO ] image_tensor , image_tensor:0 (node: image_tensor) : u8 / [N,H,W,C] / [1,300,300,3] + [ INFO ] Model outputs: + [ INFO ] detection_boxes:0 (node: DetectionOutput) : f32 / [...] / [1,1,100,7] + [Step 5/11] Resizing model to match image sizes and given batch + [ INFO ] Model batch size: 1 + [Step 6/11] Configuring input of the model + [ INFO ] Model inputs: + [ INFO ] image_tensor , image_tensor:0 (node: image_tensor) : u8 / [N,H,W,C] / [1,300,300,3] + [ INFO ] Model outputs: + [ INFO ] detection_boxes:0 (node: DetectionOutput) : f32 / [...] / [1,1,100,7] + [Step 7/11] Loading the model to the device + [ INFO ] Compile model took 127.72 ms + [Step 8/11] Querying optimal runtime parameters + [ INFO ] Model: + [ INFO ] NETWORK_NAME: frozen_inference_graph + [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 1 + [ INFO ] NUM_STREAMS: 1 + [ INFO ] AFFINITY: Affinity.CORE + [ INFO ] INFERENCE_NUM_THREADS: 14 + [ INFO ] PERF_COUNT: False + [ INFO ] INFERENCE_PRECISION_HINT: + [ INFO ] PERFORMANCE_HINT: PerformanceMode.LATENCY + [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 + [Step 9/11] Creating infer requests and preparing input tensors + [ WARNING ] No input files were given for input 'image_tensor'!. This input will be filled with random values! + [ INFO ] Fill input 'image_tensor' with random values + [Step 10/11] Measuring performance (Start inference asynchronously, 1 inference requests, limits: 60000 ms duration) + [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). + [ INFO ] First inference took 4.42 ms + [Step 11/11] Dumping statistics report + [ INFO ] Count: 15304 iterations + [ INFO ] Duration: 60005.72 ms + [ INFO ] Latency: + [ INFO ] Median: 3.87 ms + [ INFO ] Average: 3.88 ms + [ INFO ] Min: 3.49 ms + [ INFO ] Max: 5.95 ms + [ INFO ] Throughput: 255.04 FPS + + +.. code:: ipython3 + + !benchmark_app -m {model_path} -d GPU -hint latency + + +.. parsed-literal:: + + [Step 1/11] Parsing and validating input arguments + [ INFO ] Parsing input parameters + [Step 2/11] Loading OpenVINO Runtime + [ INFO ] OpenVINO: + [ INFO ] Build ................................. 2022.3.0-9052-9752fafe8eb-releases/2022/3 + [ INFO ] + [ INFO ] Device info: + [ INFO ] GPU + [ INFO ] Build ................................. 2022.3.0-9052-9752fafe8eb-releases/2022/3 + [ INFO ] + [ INFO ] + [Step 3/11] Setting device configuration + [Step 4/11] Reading model files + [ INFO ] Loading model files + [ INFO ] Read model took 14.65 ms + [ INFO ] Original model I/O parameters: + [ INFO ] Model inputs: + [ INFO ] image_tensor , image_tensor:0 (node: image_tensor) : u8 / [N,H,W,C] / [1,300,300,3] + [ INFO ] Model outputs: + [ INFO ] detection_boxes:0 (node: DetectionOutput) : f32 / [...] / [1,1,100,7] + [Step 5/11] Resizing model to match image sizes and given batch + [ INFO ] Model batch size: 1 + [Step 6/11] Configuring input of the model + [ INFO ] Model inputs: + [ INFO ] image_tensor , image_tensor:0 (node: image_tensor) : u8 / [N,H,W,C] / [1,300,300,3] + [ INFO ] Model outputs: + [ INFO ] detection_boxes:0 (node: DetectionOutput) : f32 / [...] / [1,1,100,7] + [Step 7/11] Loading the model to the device + [ INFO ] Compile model took 2254.81 ms + [Step 8/11] Querying optimal runtime parameters + [ INFO ] Model: + [ INFO ] NETWORK_NAME: frozen_inference_graph + [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 1 + [ INFO ] PERF_COUNT: False + [ INFO ] MODEL_PRIORITY: Priority.MEDIUM + [ INFO ] GPU_HOST_TASK_PRIORITY: Priority.MEDIUM + [ INFO ] GPU_QUEUE_PRIORITY: Priority.MEDIUM + [ INFO ] GPU_QUEUE_THROTTLE: Priority.MEDIUM + [ INFO ] GPU_ENABLE_LOOP_UNROLLING: True + [ INFO ] CACHE_DIR: + [ INFO ] PERFORMANCE_HINT: PerformanceMode.LATENCY + [ INFO ] COMPILATION_NUM_THREADS: 20 + [ INFO ] NUM_STREAMS: 1 + [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 + [ INFO ] INFERENCE_PRECISION_HINT: + [ INFO ] DEVICE_ID: 0 + [Step 9/11] Creating infer requests and preparing input tensors + [ WARNING ] No input files were given for input 'image_tensor'!. This input will be filled with random values! + [ INFO ] Fill input 'image_tensor' with random values + [Step 10/11] Measuring performance (Start inference asynchronously, 1 inference requests, limits: 60000 ms duration) + [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). + [ INFO ] First inference took 8.79 ms + [Step 11/11] Dumping statistics report + [ INFO ] Count: 11354 iterations + [ INFO ] Duration: 60007.21 ms + [ INFO ] Latency: + [ INFO ] Median: 4.57 ms + [ INFO ] Average: 5.16 ms + [ INFO ] Min: 3.18 ms + [ INFO ] Max: 34.87 ms + [ INFO ] Throughput: 189.21 FPS + + +CPU vs GPU with Throughput Hint `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + + +.. code:: ipython3 + + !benchmark_app -m {model_path} -d CPU -hint throughput + + +.. parsed-literal:: + + [Step 1/11] Parsing and validating input arguments + [ INFO ] Parsing input parameters + [Step 2/11] Loading OpenVINO Runtime + [ INFO ] OpenVINO: + [ INFO ] Build ................................. 2022.3.0-9052-9752fafe8eb-releases/2022/3 + [ INFO ] + [ INFO ] Device info: + [ INFO ] CPU + [ INFO ] Build ................................. 2022.3.0-9052-9752fafe8eb-releases/2022/3 + [ INFO ] + [ INFO ] + [Step 3/11] Setting device configuration + [Step 4/11] Reading model files + [ INFO ] Loading model files + [ INFO ] Read model took 29.56 ms + [ INFO ] Original model I/O parameters: + [ INFO ] Model inputs: + [ INFO ] image_tensor:0 , image_tensor (node: image_tensor) : u8 / [N,H,W,C] / [1,300,300,3] + [ INFO ] Model outputs: + [ INFO ] detection_boxes:0 (node: DetectionOutput) : f32 / [...] / [1,1,100,7] + [Step 5/11] Resizing model to match image sizes and given batch + [ INFO ] Model batch size: 1 + [Step 6/11] Configuring input of the model + [ INFO ] Model inputs: + [ INFO ] image_tensor:0 , image_tensor (node: image_tensor) : u8 / [N,H,W,C] / [1,300,300,3] + [ INFO ] Model outputs: + [ INFO ] detection_boxes:0 (node: DetectionOutput) : f32 / [...] / [1,1,100,7] + [Step 7/11] Loading the model to the device + [ INFO ] Compile model took 158.91 ms + [Step 8/11] Querying optimal runtime parameters + [ INFO ] Model: + [ INFO ] NETWORK_NAME: frozen_inference_graph + [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 5 + [ INFO ] NUM_STREAMS: 5 + [ INFO ] AFFINITY: Affinity.CORE + [ INFO ] INFERENCE_NUM_THREADS: 20 + [ INFO ] PERF_COUNT: False + [ INFO ] INFERENCE_PRECISION_HINT: + [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT + [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 + [Step 9/11] Creating infer requests and preparing input tensors + [ WARNING ] No input files were given for input 'image_tensor'!. This input will be filled with random values! + [ INFO ] Fill input 'image_tensor' with random values + [Step 10/11] Measuring performance (Start inference asynchronously, 5 inference requests, limits: 60000 ms duration) + [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). + [ INFO ] First inference took 8.15 ms + [Step 11/11] Dumping statistics report + [ INFO ] Count: 25240 iterations + [ INFO ] Duration: 60010.99 ms + [ INFO ] Latency: + [ INFO ] Median: 10.16 ms + [ INFO ] Average: 11.84 ms + [ INFO ] Min: 7.96 ms + [ INFO ] Max: 37.53 ms + [ INFO ] Throughput: 420.59 FPS + + +.. code:: ipython3 + + !benchmark_app -m {model_path} -d GPU -hint throughput + + +.. parsed-literal:: + + [Step 1/11] Parsing and validating input arguments + [ INFO ] Parsing input parameters + [Step 2/11] Loading OpenVINO Runtime + [ INFO ] OpenVINO: + [ INFO ] Build ................................. 2022.3.0-9052-9752fafe8eb-releases/2022/3 + [ INFO ] + [ INFO ] Device info: + [ INFO ] GPU + [ INFO ] Build ................................. 2022.3.0-9052-9752fafe8eb-releases/2022/3 + [ INFO ] + [ INFO ] + [Step 3/11] Setting device configuration + [Step 4/11] Reading model files + [ INFO ] Loading model files + [ INFO ] Read model took 15.45 ms + [ INFO ] Original model I/O parameters: + [ INFO ] Model inputs: + [ INFO ] image_tensor , image_tensor:0 (node: image_tensor) : u8 / [N,H,W,C] / [1,300,300,3] + [ INFO ] Model outputs: + [ INFO ] detection_boxes:0 (node: DetectionOutput) : f32 / [...] / [1,1,100,7] + [Step 5/11] Resizing model to match image sizes and given batch + [ INFO ] Model batch size: 1 + [Step 6/11] Configuring input of the model + [ INFO ] Model inputs: + [ INFO ] image_tensor , image_tensor:0 (node: image_tensor) : u8 / [N,H,W,C] / [1,300,300,3] + [ INFO ] Model outputs: + [ INFO ] detection_boxes:0 (node: DetectionOutput) : f32 / [...] / [1,1,100,7] + [Step 7/11] Loading the model to the device + [ INFO ] Compile model took 2249.04 ms + [Step 8/11] Querying optimal runtime parameters + [ INFO ] Model: + [ INFO ] NETWORK_NAME: frozen_inference_graph + [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 4 + [ INFO ] PERF_COUNT: False + [ INFO ] MODEL_PRIORITY: Priority.MEDIUM + [ INFO ] GPU_HOST_TASK_PRIORITY: Priority.MEDIUM + [ INFO ] GPU_QUEUE_PRIORITY: Priority.MEDIUM + [ INFO ] GPU_QUEUE_THROTTLE: Priority.MEDIUM + [ INFO ] GPU_ENABLE_LOOP_UNROLLING: True + [ INFO ] CACHE_DIR: + [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT + [ INFO ] COMPILATION_NUM_THREADS: 20 + [ INFO ] NUM_STREAMS: 2 + [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 + [ INFO ] INFERENCE_PRECISION_HINT: + [ INFO ] DEVICE_ID: 0 + [Step 9/11] Creating infer requests and preparing input tensors + [ WARNING ] No input files were given for input 'image_tensor'!. This input will be filled with random values! + [ INFO ] Fill input 'image_tensor' with random values + [Step 10/11] Measuring performance (Start inference asynchronously, 4 inference requests, limits: 60000 ms duration) + [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). + [ INFO ] First inference took 9.17 ms + [Step 11/11] Dumping statistics report + [ INFO ] Count: 19588 iterations + [ INFO ] Duration: 60023.47 ms + [ INFO ] Latency: + [ INFO ] Median: 11.31 ms + [ INFO ] Average: 12.15 ms + [ INFO ] Min: 9.26 ms + [ INFO ] Max: 36.04 ms + [ INFO ] Throughput: 326.34 FPS + + +Single GPU vs Multiple GPUs `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + + +.. code:: ipython3 + + !benchmark_app -m {model_path} -d GPU.1 -hint throughput + + +.. parsed-literal:: + + [Step 1/11] Parsing and validating input arguments + [ INFO ] Parsing input parameters + [Step 2/11] Loading OpenVINO Runtime + [ INFO ] OpenVINO: + [ INFO ] Build ................................. 2022.3.0-9052-9752fafe8eb-releases/2022/3 + [ INFO ] + [ INFO ] Device info: + [ INFO ] GPU + [ INFO ] Build ................................. 2022.3.0-9052-9752fafe8eb-releases/2022/3 + [ INFO ] + [ INFO ] + [Step 3/11] Setting device configuration + [ WARNING ] Device GPU.1 does not support performance hint property(-hint). + [ ERROR ] Config for device with 1 ID is not registered in GPU plugin + Traceback (most recent call last): + File "/home/adrian/repos/openvino_notebooks/venv/lib/python3.9/site-packages/openvino/tools/benchmark/main.py", line 329, in main + benchmark.set_config(config) + File "/home/adrian/repos/openvino_notebooks/venv/lib/python3.9/site-packages/openvino/tools/benchmark/benchmark.py", line 57, in set_config + self.core.set_property(device, config[device]) + RuntimeError: Config for device with 1 ID is not registered in GPU plugin + + +.. code:: ipython3 + + !benchmark_app -m {model_path} -d AUTO:GPU.1,GPU.0 -hint cumulative_throughput + + +.. parsed-literal:: + + [Step 1/11] Parsing and validating input arguments + [ INFO ] Parsing input parameters + [Step 2/11] Loading OpenVINO Runtime + [ INFO ] OpenVINO: + [ INFO ] Build ................................. 2022.3.0-9052-9752fafe8eb-releases/2022/3 + [ INFO ] + [ INFO ] Device info: + [ INFO ] AUTO + [ INFO ] Build ................................. 2022.3.0-9052-9752fafe8eb-releases/2022/3 + [ INFO ] GPU + [ INFO ] Build ................................. 2022.3.0-9052-9752fafe8eb-releases/2022/3 + [ INFO ] + [ INFO ] + [Step 3/11] Setting device configuration + [ WARNING ] Device GPU.1 does not support performance hint property(-hint). + [Step 4/11] Reading model files + [ INFO ] Loading model files + [ INFO ] Read model took 26.66 ms + [ INFO ] Original model I/O parameters: + [ INFO ] Model inputs: + [ INFO ] image_tensor , image_tensor:0 (node: image_tensor) : u8 / [N,H,W,C] / [1,300,300,3] + [ INFO ] Model outputs: + [ INFO ] detection_boxes:0 (node: DetectionOutput) : f32 / [...] / [1,1,100,7] + [Step 5/11] Resizing model to match image sizes and given batch + [ INFO ] Model batch size: 1 + [Step 6/11] Configuring input of the model + [ INFO ] Model inputs: + [ INFO ] image_tensor , image_tensor:0 (node: image_tensor) : u8 / [N,H,W,C] / [1,300,300,3] + [ INFO ] Model outputs: + [ INFO ] detection_boxes:0 (node: DetectionOutput) : f32 / [...] / [1,1,100,7] + [Step 7/11] Loading the model to the device + [ ERROR ] Config for device with 1 ID is not registered in GPU plugin + Traceback (most recent call last): + File "/home/adrian/repos/openvino_notebooks/venv/lib/python3.9/site-packages/openvino/tools/benchmark/main.py", line 414, in main + compiled_model = benchmark.core.compile_model(model, benchmark.device) + File "/home/adrian/repos/openvino_notebooks/venv/lib/python3.9/site-packages/openvino/runtime/ie_api.py", line 399, in compile_model + super().compile_model(model, device_name, {} if config is None else config), + RuntimeError: Config for device with 1 ID is not registered in GPU plugin + + +.. code:: ipython3 + + !benchmark_app -m {model_path} -d MULTI:GPU.1,GPU.0 -hint throughput + + +.. parsed-literal:: + + [Step 1/11] Parsing and validating input arguments + [ INFO ] Parsing input parameters + [Step 2/11] Loading OpenVINO Runtime + [ INFO ] OpenVINO: + [ INFO ] Build ................................. 2022.3.0-9052-9752fafe8eb-releases/2022/3 + [ INFO ] + [ INFO ] Device info: + [ INFO ] GPU + [ INFO ] Build ................................. 2022.3.0-9052-9752fafe8eb-releases/2022/3 + [ INFO ] MULTI + [ INFO ] Build ................................. 2022.3.0-9052-9752fafe8eb-releases/2022/3 + [ INFO ] + [ INFO ] + [Step 3/11] Setting device configuration + [ WARNING ] Device GPU.1 does not support performance hint property(-hint). + [Step 4/11] Reading model files + [ INFO ] Loading model files + [ INFO ] Read model took 14.84 ms + [ INFO ] Original model I/O parameters: + [ INFO ] Model inputs: + [ INFO ] image_tensor:0 , image_tensor (node: image_tensor) : u8 / [N,H,W,C] / [1,300,300,3] + [ INFO ] Model outputs: + [ INFO ] detection_boxes:0 (node: DetectionOutput) : f32 / [...] / [1,1,100,7] + [Step 5/11] Resizing model to match image sizes and given batch + [ INFO ] Model batch size: 1 + [Step 6/11] Configuring input of the model + [ INFO ] Model inputs: + [ INFO ] image_tensor:0 , image_tensor (node: image_tensor) : u8 / [N,H,W,C] / [1,300,300,3] + [ INFO ] Model outputs: + [ INFO ] detection_boxes:0 (node: DetectionOutput) : f32 / [...] / [1,1,100,7] + [Step 7/11] Loading the model to the device + [ ERROR ] Config for device with 1 ID is not registered in GPU plugin + Traceback (most recent call last): + File "/home/adrian/repos/openvino_notebooks/venv/lib/python3.9/site-packages/openvino/tools/benchmark/main.py", line 414, in main + compiled_model = benchmark.core.compile_model(model, benchmark.device) + File "/home/adrian/repos/openvino_notebooks/venv/lib/python3.9/site-packages/openvino/runtime/ie_api.py", line 399, in compile_model + super().compile_model(model, device_name, {} if config is None else config), + RuntimeError: Config for device with 1 ID is not registered in GPU plugin + + +Basic Application Using GPUs `⇑ <#top>`__ +############################################################################################################################### + + +We will now show an end-to-end object detection example using GPUs in +OpenVINO. The application compiles a model on GPU with the “THROUGHPUT” +hint, then loads a video and preprocesses every frame to convert them to +the shape expected by the model. Once the frames are loaded, it sets up +an asynchronous pipeline, performs inference and saves the detections +found in each frame. The detections are then drawn on their +corresponding frame and saved as a video, which is displayed at the end +of the application. + +Import Necessary Packages `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +.. code:: ipython3 + + import time + from pathlib import Path + + import cv2 + import numpy as np + from IPython.display import Video + from openvino.runtime import AsyncInferQueue, Core, InferRequest + + # Instantiate OpenVINO Runtime + core = Core() + core.available_devices + + + + +.. parsed-literal:: + + ['CPU', 'GPU'] + + + +Compile the Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +.. code:: ipython3 + + # Read model and compile it on GPU in THROUGHPUT mode + model = core.read_model(model=model_path) + device_name = "GPU" + compiled_model = core.compile_model(model=model, device_name=device_name, config={"PERFORMANCE_HINT": "THROUGHPUT"}) + + # Get the input and output nodes + input_layer = compiled_model.input(0) + output_layer = compiled_model.output(0) + + # Get the input size + num, height, width, channels = input_layer.shape + print('Model input shape:', num, height, width, channels) + + +.. parsed-literal:: + + Model input shape: 1 300 300 3 + + +Load and Preprocess Video Frames `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +.. code:: ipython3 + + # Load video + video_file = "https://storage.openvinotoolkit.org/repositories/openvino_notebooks/data/data/video/Coco%20Walking%20in%20Berkeley.mp4" + video = cv2.VideoCapture(video_file) + framebuf = [] + + # Go through every frame of video and resize it + print('Loading video...') + while video.isOpened(): + ret, frame = video.read() + if not ret: + print('Video loaded!') + video.release() + break + + # Preprocess frames - convert them to shape expected by model + input_frame = cv2.resize(src=frame, dsize=(width, height), interpolation=cv2.INTER_AREA) + input_frame = np.expand_dims(input_frame, axis=0) + + # Append frame to framebuffer + framebuf.append(input_frame) + + + print('Frame shape: ', framebuf[0].shape) + print('Number of frames: ', len(framebuf)) + + # Show original video file + # If the video does not display correctly inside the notebook, please open it with your favorite media player + Video(video_file) + + +.. parsed-literal:: + + Loading video... + Video loaded! + Frame shape: (1, 300, 300, 3) + Number of frames: 288 + + + + +.. raw:: html + + + + + +Define Model Output Classes `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +.. code:: ipython3 + + # Define the model's labelmap (this model uses COCO classes) + classes = [ + "background", "person", "bicycle", "car", "motorcycle", "airplane", "bus", "train", + "truck", "boat", "traffic light", "fire hydrant", "street sign", "stop sign", + "parking meter", "bench", "bird", "cat", "dog", "horse", "sheep", "cow", "elephant", + "bear", "zebra", "giraffe", "hat", "backpack", "umbrella", "shoe", "eye glasses", + "handbag", "tie", "suitcase", "frisbee", "skis", "snowboard", "sports ball", "kite", + "baseball bat", "baseball glove", "skateboard", "surfboard", "tennis racket", "bottle", + "plate", "wine glass", "cup", "fork", "knife", "spoon", "bowl", "banana", "apple", + "sandwich", "orange", "broccoli", "carrot", "hot dog", "pizza", "donut", "cake", "chair", + "couch", "potted plant", "bed", "mirror", "dining table", "window", "desk", "toilet", + "door", "tv", "laptop", "mouse", "remote", "keyboard", "cell phone", "microwave", "oven", + "toaster", "sink", "refrigerator", "blender", "book", "clock", "vase", "scissors", + "teddy bear", "hair drier", "toothbrush", "hair brush" + ] + +Set up Asynchronous Pipeline `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Callback Definition `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + + +.. code:: ipython3 + + # Define a callback function that runs every time the asynchronous pipeline completes inference on a frame + def completion_callback(infer_request: InferRequest, frame_id: int) -> None: + global frame_number + stop_time = time.time() + frame_number += 1 + + predictions = next(iter(infer_request.results.values())) + results[frame_id] = predictions[:10] # Grab first 10 predictions for this frame + + total_time = stop_time - start_time + frame_fps[frame_id] = frame_number / total_time + +Create Async Pipeline `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + + +.. code:: ipython3 + + # Create asynchronous inference queue with optimal number of infer requests + infer_queue = AsyncInferQueue(compiled_model) + infer_queue.set_callback(completion_callback) + +Perform Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +.. code:: ipython3 + + # Perform inference on every frame in the framebuffer + results = {} + frame_fps = {} + frame_number = 0 + start_time = time.time() + for i, input_frame in enumerate(framebuf): + infer_queue.start_async({0: input_frame}, i) + + infer_queue.wait_all() # Wait until all inference requests in the AsyncInferQueue are completed + stop_time = time.time() + + # Calculate total inference time and FPS + total_time = stop_time - start_time + fps = len(framebuf) / total_time + time_per_frame = 1 / fps + print(f'Total time to infer all frames: {total_time:.3f}s') + print(f'Time per frame: {time_per_frame:.6f}s ({fps:.3f} FPS)') + + +.. parsed-literal:: + + Total time to infer all frames: 1.366s + Time per frame: 0.004744s (210.774 FPS) + + +Process Results `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +.. code:: ipython3 + + # Set minimum detection threshold + min_thresh = .6 + + # Load video + video = cv2.VideoCapture(video_file) + + # Get video parameters + frame_width = int(video.get(cv2.CAP_PROP_FRAME_WIDTH)) + frame_height = int(video.get(cv2.CAP_PROP_FRAME_HEIGHT)) + fps = int(video.get(cv2.CAP_PROP_FPS)) + fourcc = int(video.get(cv2.CAP_PROP_FOURCC)) + + # Create folder and VideoWriter to save output video + Path('./output').mkdir(exist_ok=True) + output = cv2.VideoWriter('output/output.mp4', fourcc, fps, (frame_width, frame_height)) + + # Draw detection results on every frame of video and save as a new video file + while video.isOpened(): + current_frame = int(video.get(cv2.CAP_PROP_POS_FRAMES)) + ret, frame = video.read() + if not ret: + print('Video loaded!') + output.release() + video.release() + break + + # Draw info at the top left such as current fps, the devices and the performance hint being used + cv2.putText(frame, f"fps {str(round(frame_fps[current_frame], 2))}", (5, 20), cv2.FONT_ITALIC, 0.6, (0, 0, 0), 1, cv2.LINE_AA) + cv2.putText(frame, f"device {device_name}", (5, 40), cv2.FONT_ITALIC, 0.6, (0, 0, 0), 1, cv2.LINE_AA) + cv2.putText(frame, f"hint {compiled_model.get_property('PERFORMANCE_HINT').name}", (5, 60), cv2.FONT_ITALIC, 0.6, (0, 0, 0), 1, cv2.LINE_AA) + + # prediction contains [image_id, label, conf, x_min, y_min, x_max, y_max] according to model + for prediction in np.squeeze(results[current_frame]): + if prediction[2] > min_thresh: + x_min = int(prediction[3] * frame_width) + y_min = int(prediction[4] * frame_height) + x_max = int(prediction[5] * frame_width) + y_max = int(prediction[6] * frame_height) + label = classes[int(prediction[1])] + + # Draw a bounding box with its label above it + cv2.rectangle(frame, (x_min, y_min), (x_max, y_max), (0, 255, 0), 1, cv2.LINE_AA) + cv2.putText(frame, label, (x_min, y_min - 10), cv2.FONT_ITALIC, 1, (255, 0, 0), 1, cv2.LINE_AA) + + output.write(frame) + + # Show output video file + # If the video does not display correctly inside the notebook, please open it with your favorite media player + Video("output/output.mp4", width=800, embed=True) + + +.. parsed-literal:: + + Video loaded! + + + + +.. raw:: html + + + + + +Conclusion `⇑ <#top>`__ +############################################################################################################################### + + +This tutorial demonstrates how easy it is to use one or more GPUs in +OpenVINO, check their properties, and even tailor the model performance +through the different performance hints. It also provides a walk-through +of a basic object detection application that uses a GPU and displays the +detected bounding boxes. + +To read more about any of these topics, feel free to visit their +corresponding documentation: + +- `GPU + Plugin `__ +- `AUTO + Plugin `__ +- `Model + Caching `__ +- `MULTI Device + Mode `__ +- `Query Device + Properties `__ +- `Configurations for GPUs with + OpenVINO `__ +- `Benchmark Python + Tool `__ +- `Asynchronous + Inferencing `__ diff --git a/docs/notebooks/109-latency-tricks-with-output.rst b/docs/notebooks/109-latency-tricks-with-output.rst index 24202849cc3..5902b753e11 100644 --- a/docs/notebooks/109-latency-tricks-with-output.rst +++ b/docs/notebooks/109-latency-tricks-with-output.rst @@ -1,6 +1,8 @@ Performance tricks in OpenVINO for latency mode =============================================== +.. _top: + The goal of this notebook is to provide a step-by-step tutorial for improving performance for inferencing in a latency mode. Low latency is especially desired in real-time applications when the results are needed @@ -11,8 +13,8 @@ simulate a camera application that provides frames one by one. The performance tips applied in this notebook could be summarized in the following figure. Some of the steps below can be applied to any device -at any stage, e.g., “shared_memory”; some can be used only to specific -devices, e.g., “inference num threads” to CPU. As the number of +at any stage, e.g., ``shared_memory``; some can be used only to specific +devices, e.g., ``INFERENCE_NUM_THREADS`` to CPU. As the number of potential configurations is vast, we recommend looking at the steps below and then apply a trial-and-error approach. You can incorporate many hints simultaneously, like more inference threads + shared memory. @@ -38,6 +40,33 @@ optimize performance on OpenVINO IR files in run this notebook on your computer with your model to learn which of them makes sense in your case. + All the following tricks were run with OpenVINO 2022.3. Future + versions of OpenVINO may include various optimizations that may + result in different performance. + +A similar notebook focused on the throughput mode is available +`here <109-throughput-tricks.ipynb>`__. + +**Table of contents**: + +- `Data <#data>`__ +- `Model <#model>`__ +- `Hardware <#hardware>`__ +- `Helper functions <#helper-functions>`__ +- `Optimizations <#optimizations>`__ + + - `PyTorch model <#pytorch-model>`__ + - `ONNX model <#onnx-model>`__ + - `OpenVINO IR model <#openvino-ir-model>`__ + - `OpenVINO IR model on GPU <#openvino-ir-model-on-gpu>`__ + - `OpenVINO IR model + more inference threads <#openvino-ir-model-+-more-inference-threads>`__ + - `OpenVINO IR model in latency mode <#openvino-ir-model-in-latency-mode>`__ + - `OpenVINO IR model in latency mode + shared memory <#openvino-ir-model-in-latency-mode-+-shared-memory>`__ + - `Other tricks <#other-tricks>`__ + +- `Performance comparison <#performance-comparison>`__ +- `Conclusions <#conclusions>`__ + Prerequisites ------------- @@ -58,8 +87,9 @@ Prerequisites sys.path.append("../utils") import notebook_utils as utils -Data ----- +Data `⇑ <#top>`__ +############################################################################################################################### + We will use the same image of the dog sitting on a bicycle for all experiments below. The image is resized and preprocessed to fulfill the @@ -94,12 +124,13 @@ requirements of this particular object detection model. .. parsed-literal:: - + -Model ------ +Model `⇑ <#top>`__ +############################################################################################################################### + We decided to go with `YOLOv5n `__, one of the @@ -155,8 +186,9 @@ PyTorch Hub and small enough to see the difference in performance. Adding AutoShape... -Hardware --------- +Hardware `⇑ <#top>`__ +############################################################################################################################### + The code below lists the available hardware we will use in the benchmarking process. @@ -182,8 +214,9 @@ benchmarking process. CPU: Intel(R) Core(TM) i9-10920X CPU @ 3.50GHz -Helper functions ----------------- +Helper functions `⇑ <#top>`__ +############################################################################################################################### + We’re defining a benchmark model function to use for all optimized models below. It runs inference 1000 times, averages the latency time, @@ -317,15 +350,17 @@ the image. utils.show_array(output_img) -Optimizations -------------- +Optimizations `⇑ <#top>`__ +############################################################################################################################### + Below, we present the performance tricks for faster inference in the latency mode. We release resources after every benchmarking to be sure the same amount of resource is available for every experiment. -PyTorch model -~~~~~~~~~~~~~ +PyTorch model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + First, we’re benchmarking the original PyTorch model without any optimizations applied. We will treat it as our baseline. @@ -346,12 +381,13 @@ optimizations applied. We will treat it as our baseline. .. parsed-literal:: - PyTorch model on CPU. First inference time: 0.0286 seconds - PyTorch model on CPU: 0.0202 seconds per image (49.58 FPS) + PyTorch model on CPU. First inference time: 0.0227 seconds + PyTorch model on CPU: 0.0191 seconds per image (52.34 FPS) -ONNX model -~~~~~~~~~~ +ONNX model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The first optimization is exporting the PyTorch model to ONNX and running it in OpenVINO. It’s possible, thanks to the ONNX frontend. It @@ -395,18 +431,19 @@ Representation (IR) to leverage the OpenVINO Runtime. .. parsed-literal:: - ONNX model on CPU. First inference time: 0.0173 seconds - ONNX model on CPU: 0.0133 seconds per image (75.16 FPS) + ONNX model on CPU. First inference time: 0.0182 seconds + ONNX model on CPU: 0.0135 seconds per image (74.31 FPS) -OpenVINO IR model -~~~~~~~~~~~~~~~~~ +OpenVINO IR model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Let’s convert the ONNX model to OpenVINO Intermediate Representation (IR) FP16 and run it. Reducing the precision is one of the well-known methods for faster inference provided the hardware that supports lower precision, such as FP16 or even INT8. If the hardware doesn’t support -lower precisions, the model will be inferred in FP32 automatically. We +lower precision, the model will be inferred in FP32 automatically. We could also use quantization (INT8), but we should experience a little accuracy drop. That’s why we skip that step in this notebook. @@ -433,12 +470,13 @@ accuracy drop. That’s why we skip that step in this notebook. .. parsed-literal:: - OpenVINO model on CPU. First inference time: 0.0163 seconds - OpenVINO model on CPU: 0.0132 seconds per image (75.64 FPS) + OpenVINO model on CPU. First inference time: 0.0157 seconds + OpenVINO model on CPU: 0.0134 seconds per image (74.40 FPS) -OpenVINO IR model on GPU -~~~~~~~~~~~~~~~~~~~~~~~~ +OpenVINO IR model on GPU `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Usually, a GPU device is faster than a CPU, so let’s run the above model on the GPU. Please note you need to have an Intel GPU and `install @@ -461,8 +499,9 @@ execution. del ov_gpu_model # release resources -OpenVINO IR model + more inference threads -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +OpenVINO IR model + more inference threads `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + There is a possibility to add a config for any device (CPU in this case). We will increase the number of threads to an equal number of our @@ -491,11 +530,12 @@ our case. .. parsed-literal:: OpenVINO model + more threads on CPU. First inference time: 0.0151 seconds - OpenVINO model + more threads on CPU: 0.0132 seconds per image (75.68 FPS) + OpenVINO model + more threads on CPU: 0.0134 seconds per image (74.36 FPS) -OpenVINO IR model in latency mode -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +OpenVINO IR model in latency mode `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + OpenVINO offers a virtual device called `AUTO `__, @@ -520,12 +560,13 @@ devices as well. .. parsed-literal:: - OpenVINO model on AUTO. First inference time: 0.0156 seconds - OpenVINO model on AUTO: 0.0135 seconds per image (73.93 FPS) + OpenVINO model on AUTO. First inference time: 0.0159 seconds + OpenVINO model on AUTO: 0.0136 seconds per image (73.57 FPS) -OpenVINO IR model in latency mode + shared memory -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +OpenVINO IR model in latency mode + shared memory `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + OpenVINO is a C++ toolkit with Python wrappers (API). The default behavior in the Python API is copying the input to the additional buffer @@ -554,21 +595,24 @@ performance! .. parsed-literal:: - OpenVINO model + shared memory on AUTO. First inference time: 0.0144 seconds - OpenVINO model + shared memory on AUTO: 0.0053 seconds per image (187.92 FPS) + OpenVINO model + shared memory on AUTO. First inference time: 0.0139 seconds + OpenVINO model + shared memory on AUTO: 0.0054 seconds per image (184.61 FPS) -Other tricks -~~~~~~~~~~~~ +Other tricks `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -There are other tricks for performance improvement, especially -quantization and prepostprocessing. To get even more from your model, -please visit -`111-detection-quantization <../111-detection-quantization>`__ and -`118-optimize-preprocessing <../118-optimize-preprocessing>`__. -Performance comparison ----------------------- +There are other tricks for performance improvement, such as quantization +and pre-post-processing or dedicated to throughput mode. To get even +more from your model, please visit +`111-detection-quantization <../111-detection-quantization>`__, +`118-optimize-preprocessing <../118-optimize-preprocessing>`__, and +`109-throughput-tricks <109-throughput-tricks.ipynb>`__. + +Performance comparison `⇑ <#top>`__ +############################################################################################################################### + The following graphical comparison is valid for the selected model and hardware simultaneously. If you cannot see any improvement between some @@ -604,8 +648,9 @@ steps, just skip them. .. image:: 109-latency-tricks-with-output_files/109-latency-tricks-with-output_30_0.png -Conclusions ------------ +Conclusions `⇑ <#top>`__ +############################################################################################################################### + We already showed the steps needed to improve the performance of an object detection model. Even if you experience much better performance diff --git a/docs/notebooks/109-latency-tricks-with-output_files/109-latency-tricks-with-output_30_0.png b/docs/notebooks/109-latency-tricks-with-output_files/109-latency-tricks-with-output_30_0.png index d3357c378fc..741cd3a1679 100644 --- a/docs/notebooks/109-latency-tricks-with-output_files/109-latency-tricks-with-output_30_0.png +++ b/docs/notebooks/109-latency-tricks-with-output_files/109-latency-tricks-with-output_30_0.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:3683f4df707fff21d7f1e5329acbdfb9d00031ded856582d893c05a67f726d4e -size 57103 +oid sha256:9d03578132292067edf38c21858cb3b2ed61780d3ff47f1962330823f4abee26 +size 56954 diff --git a/docs/notebooks/109-latency-tricks-with-output_files/index.html b/docs/notebooks/109-latency-tricks-with-output_files/index.html index a134f651872..3fab8a55c3f 100644 --- a/docs/notebooks/109-latency-tricks-with-output_files/index.html +++ b/docs/notebooks/109-latency-tricks-with-output_files/index.html @@ -1,14 +1,14 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/109-latency-tricks-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/109-latency-tricks-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/109-latency-tricks-with-output_files/


../
-109-latency-tricks-with-output_14_0.jpg            12-Jul-2023 00:11              162715
-109-latency-tricks-with-output_17_0.jpg            12-Jul-2023 00:11              162715
-109-latency-tricks-with-output_19_0.jpg            12-Jul-2023 00:11              162756
-109-latency-tricks-with-output_23_0.jpg            12-Jul-2023 00:11              162756
-109-latency-tricks-with-output_25_0.jpg            12-Jul-2023 00:11              162756
-109-latency-tricks-with-output_27_0.jpg            12-Jul-2023 00:11              162756
-109-latency-tricks-with-output_30_0.png            12-Jul-2023 00:11               57103
-109-latency-tricks-with-output_4_0.jpg             12-Jul-2023 00:11              155828
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/109-latency-tricks-with-output_files/


../
+109-latency-tricks-with-output_14_0.jpg            16-Aug-2023 01:31              162715
+109-latency-tricks-with-output_17_0.jpg            16-Aug-2023 01:31              162715
+109-latency-tricks-with-output_19_0.jpg            16-Aug-2023 01:31              162756
+109-latency-tricks-with-output_23_0.jpg            16-Aug-2023 01:31              162756
+109-latency-tricks-with-output_25_0.jpg            16-Aug-2023 01:31              162756
+109-latency-tricks-with-output_27_0.jpg            16-Aug-2023 01:31              162756
+109-latency-tricks-with-output_30_0.png            16-Aug-2023 01:31               56954
+109-latency-tricks-with-output_4_0.jpg             16-Aug-2023 01:31              155828
 

diff --git a/docs/notebooks/109-throughput-tricks-with-output.rst b/docs/notebooks/109-throughput-tricks-with-output.rst new file mode 100644 index 00000000000..a877dab76f2 --- /dev/null +++ b/docs/notebooks/109-throughput-tricks-with-output.rst @@ -0,0 +1,717 @@ +Performance tricks in OpenVINO for throughput mode +================================================== + +.. _top: + +The goal of this notebook is to provide a step-by-step tutorial for +improving performance for inferencing in a throughput mode. High +throughput is especially desired in applications when the results are +not expected to appear as soon as possible but to lower the whole +processing time. This notebook assumes computer vision workflow and uses +`YOLOv5n `__ model. We will +simulate a video processing application that has access to all frames at +once (e.g. video editing). + +The performance tips applied in this notebook could be summarized in the +following figure. Some of the steps below can be applied to any device +at any stage, e.g., batch size; some can be used only to specific +devices, e.g., inference threads number to CPU. As the number of +potential configurations is vast, we recommend looking at the steps +below and then apply a trial-and-error approach. You can incorporate +many hints simultaneously, like more inference threads + async +processing. It should give even better performance, but we recommend +testing it anyway. + +The quantization and pre-post-processing API are not included here as +they change the precision (quantization) or processing graph +(prepostprocessor). You can find examples of how to apply them to +optimize performance on OpenVINO IR files in +`111-detection-quantization <../111-detection-quantization>`__ and +`118-optimize-preprocessing <../118-optimize-preprocessing>`__. + +|image0| + + **NOTE**: Many of the steps presented below will give you better + performance. However, some of them may not change anything if they + are strongly dependent on either the hardware or the model. Please + run this notebook on your computer with your model to learn which of + them makes sense in your case. + + All the following tricks were run with OpenVINO 2022.3. Future + versions of OpenVINO may include various optimizations that may + result in different performance. + +A similar notebook focused on the latency mode is available +`here <109-latency-tricks.ipynb>`__. + +**Table of contents**: + +- `Data <#data>`__ +- `Model <#model>`__ +- `Hardware <#hardware>`__ +- `Helper functions <#helper-functions>`__ +- `Optimizations <#optimizations>`__ + + - `PyTorch model <#pytorch-model>`__ + - `OpenVINO IR model <#openvino-ir-model>`__ + - `OpenVINO IR model + bigger batch <#openvino-ir-model-+-bigger-batch>`__ + - `OpenVINO IR model in throughput mode <#openvino-ir-model-in-throughput-mode>`__ + - `OpenVINO IR model in throughput mode on GPU <#openvino-ir-model-in-throughput-mode-on-gpu>`__ + - `OpenVINO IR model in throughput mode on AUTO <#openvino-ir-model-in-throughput-mode-on-auto>`__ + - `OpenVINO IR model in cumulative throughput mode on AUTO <#openvino-ir-model-in-cumulative-throughput-mode-on-auto>`__ + - `OpenVINO IR model in cumulative throughput mode on AUTO + asynchronous processing <#openvino-ir-model-in-cumulative-throughput-mode-on-auto-+-asynchronous-processing>`__ + - `Other tricks <#other-tricks>`__ + +- `Performance comparison <#performance-comparison>`__ +- `Conclusions <#conclusions>`__ + +Prerequisites +------------- + +.. |image0| image:: https://github.com/openvinotoolkit/openvino_notebooks/assets/4547501/e1a6e230-7c80-491a-8732-02515c556f1b + +.. code:: ipython3 + + !pip install -q seaborn ultralytics + +.. code:: ipython3 + + import sys + import time + from pathlib import Path + from typing import Any, List, Tuple + + sys.path.append("../utils") + import notebook_utils as utils + +Data `⇑ <#top>`__ +############################################################################################################################### + + +We will use the same image of the dog sitting on a bicycle copied 1000 +times to simulate the video with 1000 frames (about 33s). The image is +resized and preprocessed to fulfill the requirements of this particular +object detection model. + +.. code:: ipython3 + + import numpy as np + import cv2 + + FRAMES_NUMBER = 1024 + + IMAGE_WIDTH = 640 + IMAGE_HEIGHT = 480 + + # load image + image = utils.load_image("../data/image/coco_bike.jpg") + image = cv2.resize(image, dsize=(IMAGE_WIDTH, IMAGE_HEIGHT), interpolation=cv2.INTER_AREA) + + # preprocess it for YOLOv5 + input_image = image / 255.0 + input_image = np.transpose(input_image, axes=(2, 0, 1)) + input_image = np.expand_dims(input_image, axis=0) + + # simulate video with many frames + video_frames = np.tile(input_image, (FRAMES_NUMBER, 1, 1, 1, 1)) + + # show the image + utils.show_array(image) + + + +.. image:: 109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_4_0.jpg + + + + +.. parsed-literal:: + + + + + +Model `⇑ <#top>`__ +############################################################################################################################### + + +We decided to go with +`YOLOv5n `__, one of the +state-of-the-art object detection models, easily available through the +PyTorch Hub and small enough to see the difference in performance. + +.. code:: ipython3 + + import torch + from IPython.utils import io + + # directory for all models + base_model_dir = Path("model") + + model_name = "yolov5n" + model_path = base_model_dir / model_name + + # load YOLOv5n from PyTorch Hub + pytorch_model = torch.hub.load("ultralytics/yolov5", "custom", path=model_path, device="cpu", skip_validation=True) + # don't print full model architecture + with io.capture_output(): + pytorch_model.eval() + + +.. parsed-literal:: + + Using cache found in /opt/home/k8sworker/.cache/torch/hub/ultralytics_yolov5_master + YOLOv5 🚀 2023-4-21 Python-3.8.10 torch-1.13.1+cpu CPU + + Fusing layers... + YOLOv5n summary: 213 layers, 1867405 parameters, 0 gradients + Adding AutoShape... + + +.. parsed-literal:: + + requirements: /opt/home/k8sworker/.cache/torch/hub/requirements.txt not found, check failed. + + +Hardware `⇑ <#top>`__ +############################################################################################################################### + + +The code below lists the available hardware we will use in the +benchmarking process. + + **NOTE**: The hardware you have is probably completely different from + ours. It means you can see completely different results. + +.. code:: ipython3 + + import openvino.runtime as ov + + # initialize OpenVINO + core = ov.Core() + + # print available devices + for device in core.available_devices: + device_name = core.get_property(device, "FULL_DEVICE_NAME") + print(f"{device}: {device_name}") + + +.. parsed-literal:: + + CPU: Intel(R) Core(TM) i9-10920X CPU @ 3.50GHz + + +Helper functions `⇑ <#top>`__ +############################################################################################################################### + + +We’re defining a benchmark model function to use for all optimizations +below. It runs inference for 1000 frames and prints average frames per +second (FPS). + +.. code:: ipython3 + + from openvino.runtime import AsyncInferQueue + + + def benchmark_model(model: Any, frames: np.ndarray, async_queue: AsyncInferQueue = None, benchmark_name: str = "OpenVINO model", device_name: str = "CPU") -> float: + """ + Helper function for benchmarking the model. It measures the time and prints results. + """ + # measure the first inference separately - it may be slower as it contains also initialization + start = time.perf_counter() + model(frames[0]) + if async_queue: + async_queue.wait_all() + end = time.perf_counter() + first_infer_time = end - start + print(f"{benchmark_name} on {device_name}. First inference time: {first_infer_time :.4f} seconds") + + # benchmarking + start = time.perf_counter() + for batch in frames: + model(batch) + # wait for all threads if async processing + if async_queue: + async_queue.wait_all() + end = time.perf_counter() + + # elapsed time + infer_time = end - start + + # print second per image and FPS + mean_infer_time = infer_time / FRAMES_NUMBER + mean_fps = FRAMES_NUMBER / infer_time + print(f"{benchmark_name} on {device_name}: {mean_infer_time :.4f} seconds per image ({mean_fps :.2f} FPS)") + + return mean_fps + +The following functions aim to post-process results and draw boxes on +the image. + +.. code:: ipython3 + + # https://gist.github.com/AruniRC/7b3dadd004da04c80198557db5da4bda + classes = [ + "person", "bicycle", "car", "motorcycle", "airplane", "bus", "train", "truck", "boat", "traffic light", "fire hydrant", + "stop sign", "parking meter", "bench", "bird", "cat", "dog", "horse", "sheep", "cow", "elephant", "bear", "zebra", + "giraffe", "backpack", "umbrella", "handbag", "tie", "suitcase", "frisbee", "skis", "snowboard", "sports ball", "kite", + "baseball bat", "baseball glove", "skateboard", "surfboard", "tennis racket", "bottle", "wine glass", "cup", "fork", + "knife", "spoon", "bowl", "banana", "apple", "sandwich", "orange", "broccoli", "carrot", "hot dog", "pizza", "donut", + "cake", "chair", "couch", "potted plant", "bed", "dining table", "toilet", "tv", "laptop", "mouse", "remote", "keyboard", + "cell phone", "microwave", "oven", "oaster", "sink", "refrigerator", "book", "clock", "vase", "scissors", "teddy bear", + "hair drier", "toothbrush" + ] + + # Colors for the classes above (Rainbow Color Map). + colors = cv2.applyColorMap( + src=np.arange(0, 255, 255 / len(classes), dtype=np.float32).astype(np.uint8), + colormap=cv2.COLORMAP_RAINBOW, + ).squeeze() + + + def postprocess(detections: np.ndarray) -> List[Tuple]: + """ + Postprocess the raw results from the model. + """ + # candidates - probability > 0.25 + detections = detections[detections[..., 4] > 0.25] + + boxes = [] + labels = [] + scores = [] + for obj in detections: + xmin, ymin, ww, hh = obj[:4] + score = obj[4] + label = np.argmax(obj[5:]) + # Create a box with pixels coordinates from the box with normalized coordinates [0,1]. + boxes.append( + tuple(map(int, (xmin - ww // 2, ymin - hh // 2, ww, hh))) + ) + labels.append(int(label)) + scores.append(float(score)) + + # Apply non-maximum suppression to get rid of many overlapping entities. + # See https://paperswithcode.com/method/non-maximum-suppression + # This algorithm returns indices of objects to keep. + indices = cv2.dnn.NMSBoxes( + bboxes=boxes, scores=scores, score_threshold=0.25, nms_threshold=0.5 + ) + + # If there are no boxes. + if len(indices) == 0: + return [] + + # Filter detected objects. + return [(labels[idx], scores[idx], boxes[idx]) for idx in indices.flatten()] + + + def draw_boxes(img: np.ndarray, boxes): + """ + Draw detected boxes on the image. + """ + for label, score, box in boxes: + # Choose color for the label. + color = tuple(map(int, colors[label])) + # Draw a box. + x2 = box[0] + box[2] + y2 = box[1] + box[3] + cv2.rectangle(img=img, pt1=box[:2], pt2=(x2, y2), color=color, thickness=2) + + # Draw a label name inside the box. + cv2.putText( + img=img, + text=f"{classes[label]} {score:.2f}", + org=(box[0] + 10, box[1] + 20), + fontFace=cv2.FONT_HERSHEY_COMPLEX, + fontScale=img.shape[1] / 1200, + color=color, + thickness=1, + lineType=cv2.LINE_AA, + ) + + + def show_result(results: np.ndarray): + """ + Postprocess the raw results, draw boxes and show the image. + """ + output_img = image.copy() + + detections = postprocess(results) + draw_boxes(output_img, detections) + + utils.show_array(output_img) + +Optimizations `⇑ <#top>`__ +############################################################################################################################### + + +Below, we present the performance tricks for faster inference in the +throughput mode. We release resources after every benchmarking to be +sure the same amount of resource is available for every experiment. + +PyTorch model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +First, we’re benchmarking the original PyTorch model without any +optimizations applied. We will treat it as our baseline. + +.. code:: ipython3 + + import torch + + with torch.no_grad(): + result = pytorch_model(torch.as_tensor(video_frames[0])).detach().numpy()[0] + show_result(result) + pytorch_fps = benchmark_model(pytorch_model, frames=torch.as_tensor(video_frames).float(), benchmark_name="PyTorch model") + + + +.. image:: 109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_14_0.jpg + + +.. parsed-literal:: + + PyTorch model on CPU. First inference time: 0.0266 seconds + PyTorch model on CPU: 0.0200 seconds per image (49.99 FPS) + + +OpenVINO IR model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +The first optimization is exporting the PyTorch model to OpenVINO +Intermediate Representation (IR) FP16 and running it. Reducing the +precision is one of the well-known methods for faster inference provided +the hardware that supports lower precision, such as FP16 or even INT8. +If the hardware doesn’t support lower precision, the model will be +inferred in FP32 automatically. We could also use quantization (INT8), +but we should experience a little accuracy drop. That’s why we skip that +step in this notebook. + +.. code:: ipython3 + + from openvino.tools import mo + + onnx_path = base_model_dir / Path(f"{model_name}_{IMAGE_WIDTH}_{IMAGE_HEIGHT}").with_suffix(".onnx") + + # export PyTorch model to ONNX if it doesn't already exist + if not onnx_path.exists(): + dummy_input = torch.randn(1, 3, IMAGE_HEIGHT, IMAGE_WIDTH) + torch.onnx.export(pytorch_model, dummy_input, onnx_path) + + # convert ONNX model to IR, use FP16 + ov_model = mo.convert_model(onnx_path, compress_to_fp16=True) + +.. code:: ipython3 + + ov_cpu_model = core.compile_model(ov_model, device_name="CPU") + + result = ov_cpu_model(video_frames[0])[ov_cpu_model.output(0)][0] + show_result(result) + ov_cpu_fps = benchmark_model(model=ov_cpu_model, frames=video_frames, benchmark_name="OpenVINO model") + + del ov_cpu_model # release resources + + + +.. image:: 109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_17_0.jpg + + +.. parsed-literal:: + + OpenVINO model on CPU. First inference time: 0.0195 seconds + OpenVINO model on CPU: 0.0073 seconds per image (136.92 FPS) + + +OpenVINO IR model + bigger batch `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Batch processing often gives higher throughput as more inputs are +processed at once. To use bigger batches (than 1), we must convert the +model again, specifying a new input shape, and reshape input frames. In +our case, a batch size equal to 4 is the best choice, but optimal batch +size is very device-specific and depends on many factors, e.g., +inference precision. We recommend trying various sizes for other +hardware and model. + +.. code:: ipython3 + + batch_size = 4 + + onnx_batch_path = base_model_dir / Path(f"{model_name}_{IMAGE_WIDTH}_{IMAGE_HEIGHT}_batch_{batch_size}").with_suffix(".onnx") + + if not onnx_batch_path.exists(): + dummy_input = torch.randn(batch_size, 3, IMAGE_HEIGHT, IMAGE_WIDTH) + torch.onnx.export(pytorch_model, dummy_input, onnx_batch_path) + + # export the model with the bigger batch size + ov_batch_model = mo.convert_model(onnx_batch_path, compress_to_fp16=True) + + +.. parsed-literal:: + + /opt/home/k8sworker/.cache/torch/hub/ultralytics_yolov5_master/models/common.py:514: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + y = self.model(im, augment=augment, visualize=visualize) if augment or visualize else self.model(im) + /opt/home/k8sworker/.cache/torch/hub/ultralytics_yolov5_master/models/yolo.py:64: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + if self.dynamic or self.grid[i].shape[2:4] != x[i].shape[2:4]: + + +.. code:: ipython3 + + ov_cpu_batch_model = core.compile_model(ov_batch_model, device_name="CPU") + + batched_video_frames = video_frames.reshape([-1, batch_size, 3, IMAGE_HEIGHT, IMAGE_WIDTH]) + + result = ov_cpu_batch_model(batched_video_frames[0])[ov_cpu_batch_model.output(0)][0] + show_result(result) + ov_cpu_batch_fps = benchmark_model(model=ov_cpu_batch_model, frames=batched_video_frames, benchmark_name="OpenVINO model + bigger batch") + + del ov_cpu_batch_model # release resources + + + +.. image:: 109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_20_0.jpg + + +.. parsed-literal:: + + OpenVINO model + bigger batch on CPU. First inference time: 0.0590 seconds + OpenVINO model + bigger batch on CPU: 0.0069 seconds per image (143.96 FPS) + + +OpenVINO IR model in throughput mode `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +OpenVINO allows specifying a performance hint changing the internal +configuration of the device. There are three different hints: +``LATENCY``, ``THROUGHPUT``, and ``CUMULATIVE_THROUGHPUT``. As this +notebook is focused on the throughput mode, we will use the latter two. +The hints can be used with other devices as well. Throughput mode +implicitly triggers using the `Automatic +Batching `__ +feature, which sets the batch size to the optimal level. + +.. code:: ipython3 + + ov_cpu_through_model = core.compile_model(ov_model, device_name="CPU", config={"PERFORMANCE_HINT": "THROUGHPUT"}) + + result = ov_cpu_through_model(video_frames[0])[ov_cpu_through_model.output(0)][0] + show_result(result) + ov_cpu_through_fps = benchmark_model(model=ov_cpu_through_model, frames=video_frames, benchmark_name="OpenVINO model", device_name="CPU (THROUGHPUT)") + + del ov_cpu_through_model # release resources + + + +.. image:: 109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_22_0.jpg + + +.. parsed-literal:: + + OpenVINO model on CPU (THROUGHPUT). First inference time: 0.0226 seconds + OpenVINO model on CPU (THROUGHPUT): 0.0117 seconds per image (85.50 FPS) + + +OpenVINO IR model in throughput mode on GPU `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Usually, a GPU device provides more frames per second than a CPU, so +let’s run the above model on the GPU. Please note you need to have an +Intel GPU and `install +drivers `__ +to be able to run this step. In addition, offloading to the GPU helps +reduce CPU load and memory consumption, allowing it to be left for +routine processes. If you cannot observe a higher throughput on GPU, it +may be because the model is too light to benefit from massive parallel +execution. + +.. code:: ipython3 + + ov_gpu_fps = 0.0 + if "GPU" in core.available_devices: + # compile for GPU + ov_gpu_model = core.compile_model(ov_model, device_name="GPU", config={"PERFORMANCE_HINT": "THROUGHPUT"}) + + result = ov_gpu_model(video_frames[0])[ov_gpu_model.output(0)][0] + show_result(result) + ov_gpu_fps = benchmark_model(model=ov_gpu_model, frames=video_frames, benchmark_name="OpenVINO model", device_name="GPU (THROUGHPUT)") + + del ov_gpu_model # release resources + +OpenVINO IR model in throughput mode on AUTO `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +OpenVINO offers a virtual device called +`AUTO `__, +which can select the best device for us based on the aforementioned +performance hint. + +.. code:: ipython3 + + ov_auto_model = core.compile_model(ov_model, device_name="AUTO", config={"PERFORMANCE_HINT": "THROUGHPUT"}) + + result = ov_auto_model(video_frames[0])[ov_auto_model.output(0)][0] + show_result(result) + ov_auto_fps = benchmark_model(model=ov_auto_model, frames=video_frames, benchmark_name="OpenVINO model", device_name="AUTO (THROUGHPUT)") + + del ov_auto_model # release resources + + + +.. image:: 109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_26_0.jpg + + +.. parsed-literal:: + + OpenVINO model on AUTO (THROUGHPUT). First inference time: 0.0257 seconds + OpenVINO model on AUTO (THROUGHPUT): 0.0215 seconds per image (46.61 FPS) + + +OpenVINO IR model in cumulative throughput mode on AUTO `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +The AUTO device in throughput mode will select the best, but one +physical device to bring the highest throughput. However, if we have +more Intel devices like CPU, iGPUs, and dGPUs in one machine, we may +benefit from them all. To do so, we must use cumulative throughput to +activate all devices. + +.. code:: ipython3 + + ov_auto_cumulative_model = core.compile_model(ov_model, device_name="AUTO", config={"PERFORMANCE_HINT": "CUMULATIVE_THROUGHPUT"}) + + result = ov_auto_cumulative_model(video_frames[0])[ov_auto_cumulative_model.output(0)][0] + show_result(result) + ov_auto_cumulative_fps = benchmark_model(model=ov_auto_cumulative_model, frames=video_frames, benchmark_name="OpenVINO model", device_name="AUTO (CUMULATIVE THROUGHPUT)") + + + +.. image:: 109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_28_0.jpg + + +.. parsed-literal:: + + OpenVINO model on AUTO (CUMULATIVE THROUGHPUT). First inference time: 0.0268 seconds + OpenVINO model on AUTO (CUMULATIVE THROUGHPUT): 0.0216 seconds per image (46.25 FPS) + + +OpenVINO IR model in cumulative throughput mode on AUTO + asynchronous processing `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + +Asynchronous mode means that OpenVINO immediately returns from an +inference call and doesn’t wait for the result. It requires more +concurrent code to be written, but should offer better processing time +utilization e.g. we can run some pre- or post-processing code while +waiting for the result. Although we could use async processing directly +(start_async() function), it’s recommended to use AsyncInferQueue, which +is an easier approach to achieve the same outcome. This class +automatically spawns the pool of InferRequest objects (also called +“jobs”) and provides synchronization mechanisms to control the flow of +the pipeline. + + **NOTE**: Asynchronous processing cannot guarantee outputs to be in + the same order as inputs, so be careful in case of applications when + the order of frames matters, e.g., videos. + +.. code:: ipython3 + + from openvino.runtime import AsyncInferQueue + + + def callback(infer_request, info): + result = infer_request.get_output_tensor(0).data[0] + show_result(result) + pass + + infer_queue = AsyncInferQueue(ov_auto_cumulative_model) + infer_queue.set_callback(callback) # set callback to post-process (show) results + + infer_queue.start_async(video_frames[0]) + infer_queue.wait_all() + + # don't show output for the remaining frames + infer_queue.set_callback(lambda x, y: {}) + ov_async_model = benchmark_model(model=infer_queue.start_async, frames=video_frames, async_queue=infer_queue, benchmark_name="OpenVINO model in asynchronous processing", device_name="AUTO (CUMULATIVE THROUGHPUT)") + + del infer_queue # release resources + + + +.. image:: 109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_30_0.jpg + + +.. parsed-literal:: + + OpenVINO model in asynchronous processing on AUTO (CUMULATIVE THROUGHPUT). First inference time: 0.0239 seconds + OpenVINO model in asynchronous processing on AUTO (CUMULATIVE THROUGHPUT): 0.0041 seconds per image (245.46 FPS) + + +Other tricks `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +There are other tricks for performance improvement, such as advanced +options, quantization and pre-post-processing or dedicated to latency +mode. To get even more from your model, please visit `advanced +throughput +options `__, +`109-latency-tricks <109-latency-tricks.ipynb>`__, +`111-detection-quantization <../111-detection-quantization>`__, and +`118-optimize-preprocessing <../118-optimize-preprocessing>`__. + +Performance comparison `⇑ <#top>`__ +############################################################################################################################### + + +The following graphical comparison is valid for the selected model and +hardware simultaneously. If you cannot see any improvement between some +steps, just skip them. + +.. code:: ipython3 + + %matplotlib inline + +.. code:: ipython3 + + from matplotlib import pyplot as plt + + labels = ["PyTorch model", "OpenVINO IR model", "OpenVINO IR model + bigger batch", "OpenVINO IR model in throughput mode", "OpenVINO IR model in throughput mode on GPU", + "OpenVINO IR model in throughput mode on AUTO", "OpenVINO IR model in cumulative throughput mode on AUTO", "OpenVINO IR model in cumulative throughput mode on AUTO + asynchronous processing"] + + fps = [pytorch_fps, ov_cpu_fps, ov_cpu_batch_fps, ov_cpu_through_fps, ov_gpu_fps, ov_auto_fps, ov_auto_cumulative_fps, ov_async_model] + + bar_colors = colors[::10] / 255.0 + + fig, ax = plt.subplots(figsize=(16, 8)) + ax.bar(labels, fps, color=bar_colors) + + ax.set_ylabel("Throughput [FPS]") + ax.set_title("Performance difference") + + plt.xticks(rotation='vertical') + plt.show() + + + +.. image:: 109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_33_0.png + + +Conclusions `⇑ <#top>`__ +############################################################################################################################### + + +We already showed the steps needed to improve the throughput of an +object detection model. Even if you experience much better performance +after running this notebook, please note this may not be valid for every +hardware or every model. For the most accurate results, please use +``benchmark_app`` `command-line +tool `__. +Note that ``benchmark_app`` cannot measure the impact of some tricks +above. diff --git a/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_14_0.jpg b/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_14_0.jpg new file mode 100644 index 00000000000..ab32fedd0f3 --- /dev/null +++ b/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_14_0.jpg @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:84e4f91c248768c2ea746240e307041396099f0d52fdb89b0179fa72e353894a +size 162715 diff --git a/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_17_0.jpg b/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_17_0.jpg new file mode 100644 index 00000000000..3241ad3c3f5 --- /dev/null +++ b/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_17_0.jpg @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:fcd7da923bb0f72430eaf7b4770175320f1f3219aaca2d460c54fa9ef07e51c2 +size 162756 diff --git a/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_20_0.jpg b/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_20_0.jpg new file mode 100644 index 00000000000..3241ad3c3f5 --- /dev/null +++ b/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_20_0.jpg @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:fcd7da923bb0f72430eaf7b4770175320f1f3219aaca2d460c54fa9ef07e51c2 +size 162756 diff --git a/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_22_0.jpg b/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_22_0.jpg new file mode 100644 index 00000000000..3241ad3c3f5 --- /dev/null +++ b/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_22_0.jpg @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:fcd7da923bb0f72430eaf7b4770175320f1f3219aaca2d460c54fa9ef07e51c2 +size 162756 diff --git a/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_26_0.jpg b/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_26_0.jpg new file mode 100644 index 00000000000..3241ad3c3f5 --- /dev/null +++ b/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_26_0.jpg @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:fcd7da923bb0f72430eaf7b4770175320f1f3219aaca2d460c54fa9ef07e51c2 +size 162756 diff --git a/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_28_0.jpg b/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_28_0.jpg new file mode 100644 index 00000000000..3241ad3c3f5 --- /dev/null +++ b/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_28_0.jpg @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:fcd7da923bb0f72430eaf7b4770175320f1f3219aaca2d460c54fa9ef07e51c2 +size 162756 diff --git a/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_30_0.jpg b/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_30_0.jpg new file mode 100644 index 00000000000..3241ad3c3f5 --- /dev/null +++ b/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_30_0.jpg @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:fcd7da923bb0f72430eaf7b4770175320f1f3219aaca2d460c54fa9ef07e51c2 +size 162756 diff --git a/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_33_0.png b/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_33_0.png new file mode 100644 index 00000000000..ef164610208 --- /dev/null +++ b/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_33_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:9547ead08b2efa00ac1380ab53f7cdb22f416ea13ecf0d9eb556703e9fcdde0d +size 77855 diff --git a/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_4_0.jpg b/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_4_0.jpg new file mode 100644 index 00000000000..510b092f676 --- /dev/null +++ b/docs/notebooks/109-throughput-tricks-with-output_files/109-throughput-tricks-with-output_4_0.jpg @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:41c502fdff24ada81c63ccfca7d9153ea368b1eb3330caa03afa3281c35e4484 +size 155828 diff --git a/docs/notebooks/109-throughput-tricks-with-output_files/index.html b/docs/notebooks/109-throughput-tricks-with-output_files/index.html new file mode 100644 index 00000000000..e5167f73bef --- /dev/null +++ b/docs/notebooks/109-throughput-tricks-with-output_files/index.html @@ -0,0 +1,15 @@ + +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/109-throughput-tricks-with-output_files/ + +

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/109-throughput-tricks-with-output_files/


../
+109-throughput-tricks-with-output_14_0.jpg         16-Aug-2023 01:31              162715
+109-throughput-tricks-with-output_17_0.jpg         16-Aug-2023 01:31              162756
+109-throughput-tricks-with-output_20_0.jpg         16-Aug-2023 01:31              162756
+109-throughput-tricks-with-output_22_0.jpg         16-Aug-2023 01:31              162756
+109-throughput-tricks-with-output_26_0.jpg         16-Aug-2023 01:31              162756
+109-throughput-tricks-with-output_28_0.jpg         16-Aug-2023 01:31              162756
+109-throughput-tricks-with-output_30_0.jpg         16-Aug-2023 01:31              162756
+109-throughput-tricks-with-output_33_0.png         16-Aug-2023 01:31               77855
+109-throughput-tricks-with-output_4_0.jpg          16-Aug-2023 01:31              155828
+

+ diff --git a/docs/notebooks/110-ct-scan-live-inference-with-output.rst b/docs/notebooks/110-ct-scan-live-inference-with-output.rst index 60c039ee402..6115989b4bb 100644 --- a/docs/notebooks/110-ct-scan-live-inference-with-output.rst +++ b/docs/notebooks/110-ct-scan-live-inference-with-output.rst @@ -1,6 +1,8 @@ Live Inference and Benchmark CT-scan Data with OpenVINO™ ======================================================== +.. _top: + Kidney Segmentation with PyTorch Lightning and OpenVINO™ - Part 4 ----------------------------------------------------------------- @@ -16,25 +18,39 @@ live inference with async API and MULTI plugin in OpenVINO. This notebook needs a quantized OpenVINO IR model and images from the `KiTS-19 `__ dataset, converted to -2D images. (To learn how the model is quantized, see the `Convert and -Quantize a UNet Model and Show Live -Inference <110-ct-segmentation-quantize-nncf.ipynb>`__ tutorial.) +2D images. (To learn how the model is quantized, see the +`Convert and Quantize a UNet Model and Show Live Inference <110-ct-segmentation-quantize-nncf-with-output.html>`__ tutorial.) This notebook provides a pre-trained model, trained for 20 epochs with the full KiTS-19 frames dataset, which has an F1 score on the validation -set of 0.9. The training code is available in the `PyTorch Monai -Training <110-ct-segmentation-quantize-with-output.html>`__ +set of 0.9. The training code is available in the +`PyTorch MONAI Training <110-ct-segmentation-quantize-with-output.html>`__ notebook. For demonstration purposes, this tutorial will download one converted CT -scan to use for inference. +scan to use for inference. + +**Table of contents**: + +- `Imports <#imports>`__ +- `Settings <#settings>`__ +- `Benchmark Model Performance <#benchmark-model-performance>`__ +- `Download and Prepare Data <#download-and-prepare-data>`__ +- `Show Live Inference <#show-live-inference>`__ + + - `Load Model and List of Image Files <#load-model-and-list-of-image-files>`__ + - `Prepare images <#prepare-images>`__ + - `Specify device <#specify-device>`__ + - `Setting callback function <#setting-callback-function>`__ + - `Create asynchronous inference queue and perform it <#create-asynchronous-inference-queue-and-perform-it>`__ .. code:: ipython3 !pip install -q "monai>=0.9.1,<1.0.0" -Imports -------- +Imports `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -52,8 +68,9 @@ Imports sys.path.append("../utils") from notebook_utils import download_file -Settings --------- +Settings `⇑ <#top>`__ +############################################################################################################################### + To use the pre-trained models, set ``IR_PATH`` to ``"pretrained_model/unet44.xml"`` and ``COMPRESSED_MODEL_PATH`` to @@ -90,11 +107,11 @@ trained or optimized yourself, adjust the model paths. pretrained_model/quantized_unet_kits19.bin: 0%| | 0.00/1.90M [00:00`__ +############################################################################################################################### -To measure the inference performance of the IR model, use `Benchmark -Tool `__ +To measure the inference performance of the IR model, use +`Benchmark Tool `__ - an inference performance measurement tool in OpenVINO. Benchmark tool is a command-line application that can be run in the notebook with ``! benchmark_app`` or ``%sx benchmark_app`` commands. @@ -157,7 +174,7 @@ is a command-line application that can be run in the notebook with [ WARNING ] Performance hint was not explicitly specified in command line. Device(AUTO) performance hint will be set to PerformanceMode.LATENCY. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 13.92 ms + [ INFO ] Read model took 13.69 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] input.1 (node: input.1) : f32 / [...] / [1,1,512,512] @@ -171,7 +188,7 @@ is a command-line application that can be run in the notebook with [ INFO ] Model outputs: [ INFO ] 153 (node: 153) : f32 / [...] / [1,1,512,512] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 181.67 ms + [ INFO ] Compile model took 181.66 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: [ INFO ] PERFORMANCE_HINT: PerformanceMode.LATENCY @@ -200,21 +217,22 @@ is a command-line application that can be run in the notebook with [ INFO ] Fill input 'input.1' with random values [Step 10/11] Measuring performance (Start inference synchronously, limits: 15000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 27.22 ms + [ INFO ] First inference took 26.05 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 1431 iterations - [ INFO ] Duration: 15005.09 ms + [ INFO ] Count: 1424 iterations + [ INFO ] Duration: 15004.96 ms [ INFO ] Latency: - [ INFO ] Median: 10.25 ms - [ INFO ] Average: 10.29 ms - [ INFO ] Min: 9.99 ms - [ INFO ] Max: 15.67 ms - [ INFO ] Throughput: 97.57 FPS + [ INFO ] Median: 10.29 ms + [ INFO ] Average: 10.35 ms + [ INFO ] Min: 10.14 ms + [ INFO ] Max: 14.72 ms + [ INFO ] Throughput: 97.14 FPS -Download and Prepare Data -------------------------- +Download and Prepare Data `⇑ <#top>`__ +############################################################################################################################### + Download one validation video for live inference. @@ -260,8 +278,9 @@ downloaded and extracted in the next cell. Downloaded and extracted data for case_00117 -Show Live Inference -------------------- +Show Live Inference `⇑ <#top>`__ +############################################################################################################################### + To show live inference on the model in the notebook, use the asynchronous processing feature of OpenVINO Runtime. @@ -271,28 +290,31 @@ If you use a GPU device, with ``device="GPU"`` or card, model loading will be slow the first time you run this code. The model will be cached, so after the first time model loading will be faster. For more information on OpenVINO Runtime, including Model -Caching, refer to the `OpenVINO API -tutorial <002-openvino-api-with-output.html>`__. +Caching, refer to the `OpenVINO API tutorial <002-openvino-api-with-output.html>`__. -We will use -`AsyncInferQueue `__ +We will use `AsyncInferQueue `__ to perform asynchronous inference. It can be instantiated with compiled model and a number of jobs - parallel execution threads. If you don’t pass a number of jobs or pass ``0``, then OpenVINO will pick the optimal number based on your device and heuristics. After acquiring the inference queue, there are two jobs to do: -- Preprocess the data and push it to the inference queue. The preprocessing steps will remain the same. -- Tell the inference queue what to do with the model output after the inference is finished. It is represented by the ``callback`` python function that takes an inference result and data that we passed to the inference queue along with the prepared input data +- Preprocess the data and push it to the inference queue. The + preprocessing steps will remain the same. +- Tell the inference queue what to do with the model output after the + inference is finished. It is represented by the ``callback`` python + function that takes an inference result and data that we passed to + the inference queue along with the prepared input data Everything else will be handled by the ``AsyncInferQueue`` instance. -Load Model and List of Image Files -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Load Model and List of Image Files `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Load the segmentation model to OpenVINO Runtime with -``SegmentationModel``, based on the Model API from `Open Model -Zoo `__. This model +``SegmentationModel``, based on the Model API from +`Open Model Zoo `__. This model implementation includes pre and post processing for the model. For ``SegmentationModel`` this includes the code to create an overlay of the segmentation mask on the original image/frame. Uncomment the next cell @@ -314,8 +336,9 @@ to see the implementation. case_00117, 69 images -Preapre images -~~~~~~~~~~~~~~ +Prepare images `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Use the ``reader = LoadImage()`` function to read the images in the same way as in the @@ -335,8 +358,9 @@ tutorial. framebuf.append(image) next_frame_id += 1 -Specify device -~~~~~~~~~~~~~~ +Specify device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -351,14 +375,15 @@ Specify device -Setting callback function -~~~~~~~~~~~~~~~~~~~~~~~~~ +Setting callback function `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + When ``callback`` is set, any job that ends the inference, calls the Python function. The ``callback`` function must have two arguments: one is the request that calls the ``callback``, which provides the -InferRequest API; the other is called “userdata”, which provides the -possibility of passing runtime values. +``InferRequest`` API; the other is called ``userdata``, which provides +the possibility of passing runtime values. The ``callback`` function will show the results of inference. @@ -387,8 +412,9 @@ The ``callback`` function will show the results of inference. display.clear_output(wait=True) display.display(i) -Create asynchronous inference queue and perform it -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Create asynchronous inference queue and perform it `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -433,6 +459,6 @@ Create asynchronous inference queue and perform it .. parsed-literal:: Loaded model to Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') in 0.18 seconds. - Total time to infer all frames: 3.416s - Time per frame: 0.050229s (19.909 FPS) + Total time to infer all frames: 3.401s + Time per frame: 0.050022s (19.991 FPS) diff --git a/docs/notebooks/110-ct-scan-live-inference-with-output_files/index.html b/docs/notebooks/110-ct-scan-live-inference-with-output_files/index.html index 8e421c6cf63..bc2298c536d 100644 --- a/docs/notebooks/110-ct-scan-live-inference-with-output_files/index.html +++ b/docs/notebooks/110-ct-scan-live-inference-with-output_files/index.html @@ -1,7 +1,7 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/110-ct-scan-live-inference-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/110-ct-scan-live-inference-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/110-ct-scan-live-inference-with-output_files/


../
-110-ct-scan-live-inference-with-output_21_0.png    12-Jul-2023 00:11               48780
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/110-ct-scan-live-inference-with-output_files/


../
+110-ct-scan-live-inference-with-output_21_0.png    16-Aug-2023 01:31               48780
 

diff --git a/docs/notebooks/110-ct-segmentation-quantize-nncf-with-output.rst b/docs/notebooks/110-ct-segmentation-quantize-nncf-with-output.rst index 92f851d6e76..01c823c0cb2 100644 --- a/docs/notebooks/110-ct-segmentation-quantize-nncf-with-output.rst +++ b/docs/notebooks/110-ct-segmentation-quantize-nncf-with-output.rst @@ -1,6 +1,8 @@ Quantize a Segmentation Model and Show Live Inference ===================================================== +.. _top: + Kidney Segmentation with PyTorch Lightning and OpenVINO™ - Part 3 ----------------------------------------------------------------- @@ -13,10 +15,8 @@ scratch; the data is from This third tutorial in the series shows how to: -- Convert an Original model to OpenVINO IR with `Model - Optimizer `__, - using `Model Optimizer Python - API `__ +- Convert an Original model to OpenVINO IR with `model conversion + API `__ - Quantize a PyTorch model with NNCF - Evaluate the F1 score metric of the original model and the quantized model @@ -48,19 +48,46 @@ NNCF for PyTorch models requires a C++ compiler. On Windows, install 2019 `__. During installation, choose Desktop development with C++ in the Workloads tab. On macOS, run ``xcode-select –install`` from a Terminal. -On Linux, install gcc. +On Linux, install ``gcc``. Running this notebook with the full dataset will take a long time. For demonstration purposes, this tutorial will download one converted CT scan and use that scan for quantization and inference. For production -purposes, use a representative dataset for quantizing the model. +purposes, use a representative dataset for quantizing the model. + +**Table of contents**: + +- `Imports <#imports>`__ +- `Settings <#settings>`__ +- `Load PyTorch Model <#load-pytorch-model>`__ +- `Download CT-scan Data <#download-ct-scan-data>`__ +- `Configuration <#configuration>`__ + + - `Dataset <#dataset>`__ + - `Metric <#metric>`__ + +- `Quantization <#quantization>`__ +- `Compare FP32 and INT8 Model <#compare-fp32-and-int8-model>`__ + + - `Compare File Size <#compare-file-size>`__ + - `Compare Metrics for the original model and the quantized model to be sure that there no degradation. <#compare-metrics-for-the-original-model-and-the-quantized-model-to-be-sure-that-there-no-degradation>`__ + - `Compare Performance of the FP32 IR Model and Quantized Models <#compare-performance-of-the-fp32-ir-model-and-quantized-models>`__ + - `Visually Compare Inference Results <#visually-compare-inference-results>`__ + +- `Show Live Inference <#show-live-inference>`__ + + - `Load Model and List of Image Files <#load-model-and-list-of-image-files>`__ + - `Show Inference <#show-inference>`__ + +- `References <#references>`__ .. code:: ipython3 !pip install -q "monai>=0.9.1,<1.0.0" "torchmetrics>=0.11.0" -Imports -------- +Imports `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -146,10 +173,10 @@ Imports .. parsed-literal:: - 2023-07-11 22:37:38.702892: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. - 2023-07-11 22:37:38.737776: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + 2023-08-15 22:41:33.627938: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-08-15 22:41:33.662730: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. - 2023-07-11 22:37:39.282688: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + 2023-08-15 22:41:34.189615: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT .. parsed-literal:: @@ -157,8 +184,9 @@ Imports INFO:nncf:NNCF initialized successfully. Supported frameworks detected: torch, tensorflow, onnx, openvino -Settings --------- +Settings `⇑ <#top>`__ +############################################################################################################################### + By default, this notebook will download one CT scan from the KITS19 dataset that will be used for quantization. To use the full dataset, set @@ -173,16 +201,16 @@ dataset that will be used for quantization. To use the full dataset, set MODEL_DIR = Path("model") MODEL_DIR.mkdir(exist_ok=True) -Load PyTorch Model ------------------- +Load PyTorch Model `⇑ <#top>`__ +############################################################################################################################### + Download the pre-trained model weights, load the PyTorch model and the ``state_dict`` that was saved after training. The model used in this -notebook is a -`BasicUnet `__ +notebook is a `BasicUNet `__ model from `MONAI `__. We provide a pre-trained -checkpoint. To see how this model performs, check out the `training -notebook `__. +checkpoint. To see how this model performs, check out the +`training notebook `__. .. code:: ipython3 @@ -219,8 +247,9 @@ notebook `__. -Download CT-scan Data ---------------------- +Download CT-scan Data `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -245,19 +274,20 @@ Download CT-scan Data Data for case_00117 exists -Configuration -------------- +Configuration `⇑ <#top>`__ +############################################################################################################################### -Dataset -~~~~~~~ -The KitsDataset class in the next cell expects images and masks in the -*basedir* directory, in a folder per patient. It is a simplified version -of the DataSet class in the `training -notebook `__. +Dataset `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +The ``KitsDataset`` class in the next cell expects images and masks in +the *``basedir``* directory, in a folder per patient. It is a simplified +version of the Dataset class in the `training notebook `__. Images are loaded with MONAI’s -```LoadImage`` `__, +`LoadImage `__, to align with the image loading method in the training notebook. This method rotates and flips the images. We define a ``rotate_and_flip`` method to display the images in the expected orientation: @@ -346,8 +376,8 @@ kidney pixels to verify that the annotations look correct: .. image:: 110-ct-segmentation-quantize-nncf-with-output_files/110-ct-segmentation-quantize-nncf-with-output_15_1.png -Metric -~~~~~~ +Metric `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ Define a metric to determine the performance of the model. @@ -385,8 +415,9 @@ library. metric.update(label.flatten(), prediction.flatten()) return metric.compute() -Quantization ------------- +Quantization `⇑ <#top>`__ +############################################################################################################################### + Before quantizing the model, we compute the F1 score on the ``FP32`` model, for comparison: @@ -416,26 +447,27 @@ this notebook. .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/monai/networks/nets/basic_unet.py:179: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/monai/networks/nets/basic_unet.py:179: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if x_e.shape[-i - 1] != x_0.shape[-i - 1]: `NNCF `__ provides a suite of advanced algorithms for Neural Networks inference optimization in -OpenVINO with minimal accuracy drop. +OpenVINO with minimal accuracy drop. -.. note:: - - NNCF Post-training Quantization is available in OpenVINO 2023.0 release. + **Note**: NNCF Post-training Quantization is available in OpenVINO + 2023.0 release. Create a quantized model from the pre-trained ``FP32`` model and the calibration dataset. The optimization process contains the following steps: -1. Create a Dataset for quantization. -2. Run ``nncf.quantize`` for getting an optimized model. -3. Export the quantized model to ONNX and then convert to OpenVINO IR model. -4. Serialize the INT8 model using ``openvino.runtime.serialize`` function for benchmarking. +:: + + 1. Create a Dataset for quantization. + 2. Run `nncf.quantize` for getting an optimized model. + 3. Export the quantized model to ONNX and then convert to OpenVINO IR model. + 4. Serialize the INT8 model using `openvino.runtime.serialize` function for benchmarking. .. code:: ipython3 @@ -479,28 +511,29 @@ model and save it. .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/nncf/torch/quantization/layers.py:338: TracerWarning: Converting a tensor to a Python number might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/nncf/torch/quantization/layers.py:338: TracerWarning: Converting a tensor to a Python number might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! return self._level_low.item() - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/nncf/torch/quantization/layers.py:346: TracerWarning: Converting a tensor to a Python number might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/nncf/torch/quantization/layers.py:346: TracerWarning: Converting a tensor to a Python number might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! return self._level_high.item() - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/monai/networks/nets/basic_unet.py:179: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/monai/networks/nets/basic_unet.py:179: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if x_e.shape[-i - 1] != x_0.shape[-i - 1]: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/nncf/torch/quantization/quantize_functions.py:140: FutureWarning: 'torch.onnx._patch_torch._graph_op' is deprecated in version 1.13 and will be removed in version 1.14. Please note 'g.op()' is to be removed from torch.Graph. Please open a GitHub issue if you need this functionality.. + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/nncf/torch/quantization/quantize_functions.py:140: FutureWarning: 'torch.onnx._patch_torch._graph_op' is deprecated in version 1.13 and will be removed in version 1.14. Please note 'g.op()' is to be removed from torch.Graph. Please open a GitHub issue if you need this functionality.. output = g.op( This notebook demonstrates post-training quantization with NNCF. NNCF also supports quantization-aware training, and other algorithms -than quantization. See the `NNCF -documentation `__ in the NNCF +than quantization. See the `NNCF documentation `__ in the NNCF repository for more information. -Compare FP32 and INT8 Model ---------------------------- +Compare FP32 and INT8 Model `⇑ <#top>`__ +############################################################################################################################### + + +Compare File Size `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Compare File Size -~~~~~~~~~~~~~~~~~ .. code:: ipython3 @@ -517,8 +550,8 @@ Compare File Size INT8 model size: 1953.49 KB -Compare Metrics for the original model and the quantized model to be sure that there no degradation. -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Compare Metrics for the original model and the quantized model to be sure that there no degradation. `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ .. code:: ipython3 @@ -537,12 +570,11 @@ Compare Metrics for the original model and the quantized model to be sure that t INT8 F1: 0.999 -Compare Performance of the FP32 IR Model and Quantized Models -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Compare Performance of the FP32 IR Model and Quantized Models `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ To measure the inference performance of the ``FP32`` and ``INT8`` -models, we use `Benchmark -Tool `__ +models, we use `Benchmark Tool `__ - OpenVINO’s inference performance measurement tool. Benchmark tool is a command line application, part of OpenVINO development tools, that can be run in the notebook with ``! benchmark_app`` or @@ -586,7 +618,7 @@ be run in the notebook with ``! benchmark_app`` or [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.LATENCY. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 22.47 ms + [ INFO ] Read model took 21.45 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] 1 , x (node: Parameter_2) : f32 / [...] / [1,1,512,512] @@ -600,7 +632,7 @@ be run in the notebook with ``! benchmark_app`` or [ INFO ] Model outputs: [ INFO ] 169 (node: aten::_convolution_861) : f32 / [...] / [1,1,512,512] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 89.66 ms + [ INFO ] Compile model took 87.17 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: [ INFO ] NETWORK_NAME: Model0 @@ -622,17 +654,17 @@ be run in the notebook with ``! benchmark_app`` or [ INFO ] Fill input '1' with random values [Step 10/11] Measuring performance (Start inference synchronously, limits: 15000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 56.61 ms + [ INFO ] First inference took 55.25 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 427 iterations - [ INFO ] Duration: 15001.35 ms + [ INFO ] Count: 425 iterations + [ INFO ] Duration: 15001.92 ms [ INFO ] Latency: - [ INFO ] Median: 34.89 ms - [ INFO ] Average: 34.92 ms - [ INFO ] Min: 34.57 ms - [ INFO ] Max: 39.09 ms - [ INFO ] Throughput: 28.66 FPS + [ INFO ] Median: 35.01 ms + [ INFO ] Average: 35.08 ms + [ INFO ] Min: 34.54 ms + [ INFO ] Max: 37.24 ms + [ INFO ] Throughput: 28.56 FPS .. code:: ipython3 @@ -658,7 +690,7 @@ be run in the notebook with ``! benchmark_app`` or [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.LATENCY. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 32.41 ms + [ INFO ] Read model took 33.80 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] x.1 (node: x.1) : f32 / [...] / [1,1,512,512] @@ -672,7 +704,7 @@ be run in the notebook with ``! benchmark_app`` or [ INFO ] Model outputs: [ INFO ] 578 (node: 578) : f32 / [...] / [1,1,512,512] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 144.74 ms + [ INFO ] Compile model took 144.48 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: [ INFO ] NETWORK_NAME: torch_jit @@ -694,21 +726,22 @@ be run in the notebook with ``! benchmark_app`` or [ INFO ] Fill input 'x.1' with random values [Step 10/11] Measuring performance (Start inference synchronously, limits: 15000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 32.29 ms + [ INFO ] First inference took 30.85 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 985 iterations - [ INFO ] Duration: 15013.03 ms + [ INFO ] Count: 973 iterations + [ INFO ] Duration: 15001.74 ms [ INFO ] Latency: - [ INFO ] Median: 14.99 ms - [ INFO ] Average: 15.04 ms - [ INFO ] Min: 14.76 ms - [ INFO ] Max: 17.68 ms - [ INFO ] Throughput: 66.70 FPS + [ INFO ] Median: 15.17 ms + [ INFO ] Average: 15.21 ms + [ INFO ] Min: 14.84 ms + [ INFO ] Max: 17.66 ms + [ INFO ] Throughput: 65.90 FPS -Visually Compare Inference Results -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Visually Compare Inference Results `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Visualize the results of the model on four slices of the validation set. Compare the results of the ``FP32`` IR model with the results of the @@ -787,24 +820,23 @@ seed is displayed to enable reproducing specific runs of this cell. .. parsed-literal:: - Visualizing results with seed 1689107962 + Visualizing results with seed 1692132195 .. image:: 110-ct-segmentation-quantize-nncf-with-output_files/110-ct-segmentation-quantize-nncf-with-output_37_1.png -Show Live Inference -------------------- +Show Live Inference `⇑ <#top>`__ +############################################################################################################################### + To show live inference on the model in the notebook, we will use the asynchronous processing feature of OpenVINO. -We use the ``show_live_inference`` function from `Notebook -Utils `__ to show live inference. This -function uses `Open Model -Zoo `__\ ’s -AsyncPipeline and Model API to perform asynchronous inference. After +We use the ``show_live_inference`` function from `Notebook Utils `__ to show live inference. This +function uses `Open Model Zoo `__ +Async Pipeline and Model API to perform asynchronous inference. After inference on the specified CT scan has completed, the total time and throughput (fps), including preprocessing and displaying, will be printed. @@ -812,12 +844,12 @@ printed. **NOTE**: If you experience flickering on Firefox, consider using Chrome or Edge to run this notebook. -Load Model and List of Image Files -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Load Model and List of Image Files `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + We load the segmentation model to OpenVINO Runtime with -``SegmentationModel``, based on the `Open Model -Zoo `__ Model API. +``SegmentationModel``, based on the `Open Model Zoo `__ Model API. This model implementation includes pre and post processing for the model. For ``SegmentationModel``, this includes the code to create an overlay of the segmentation mask on the original image/frame. @@ -839,8 +871,9 @@ overlay of the segmentation mask on the original image/frame. case_00117, 69 images -Show Inference -~~~~~~~~~~~~~~ +Show Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + In the next cell, we run the ``show_live_inference`` function, which loads the ``segmentation_model`` to the specified ``device`` (using @@ -864,27 +897,24 @@ performs inference, and displays the results on the frames loaded in .. parsed-literal:: - Loaded model to CPU in 0.13 seconds. - Total time for 68 frames: 3.28 seconds, fps:21.05 + Loaded model to CPU in 0.12 seconds. + Total time for 68 frames: 3.35 seconds, fps:20.60 -References ----------- +References `⇑ <#top>`__ +############################################################################################################################### -**OpenVINO** - `NNCF -Repository `__ - `Neural -Network Compression Framework for fast model -inference `__ - `OpenVINO API -Tutorial <002-openvino-api-with-output.html>`__ - `OpenVINO -PyPI (pip install -openvino-dev) `__ -**Kits19 Data** - `Kits19 Challenge -Homepage `__ - `Kits19 Github -Repository `__ - `The KiTS19 -Challenge Data: 300 Kidney Tumor Cases with Clinical Context, CT -Semantic Segmentations, and Surgical -Outcomes `__ - `The state of the art -in kidney and kidney tumor segmentation in contrast-enhanced CT imaging: -Results of the KiTS19 -challenge `__ +**OpenVINO** + +- `NNCF Repository `__ +- `Neural Network Compression Framework for fast model inference `__ +- `OpenVINO API Tutorial <002-openvino-api-with-output.html>`__ +- `OpenVINO PyPI (pip install openvino-dev) `__ + +**Kits19 Data** + +- `Kits19 Challenge Homepage `__ +- `Kits19 GitHub Repository `__ +- `The KiTS19 Challenge Data: 300 Kidney Tumor Cases with Clinical Context, CT Semantic Segmentations, and Surgical Outcomes `__ +- `The state of the art in kidney and kidney tumor segmentation in contrast-enhanced CT imaging: Results of the KiTS19 challenge `__ diff --git a/docs/notebooks/110-ct-segmentation-quantize-nncf-with-output_files/110-ct-segmentation-quantize-nncf-with-output_37_1.png b/docs/notebooks/110-ct-segmentation-quantize-nncf-with-output_files/110-ct-segmentation-quantize-nncf-with-output_37_1.png index 789d1e38698..2ed816a7039 100644 --- a/docs/notebooks/110-ct-segmentation-quantize-nncf-with-output_files/110-ct-segmentation-quantize-nncf-with-output_37_1.png +++ b/docs/notebooks/110-ct-segmentation-quantize-nncf-with-output_files/110-ct-segmentation-quantize-nncf-with-output_37_1.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:960b979d5c041593481e188c2f1f14c72c314c8b2338c263ad450e6ad52e3600 -size 385435 +oid sha256:aee7d177b413f4c198de3e7c40f1624d59abbdabe61e8ec20798b4649ac46f43 +size 383352 diff --git a/docs/notebooks/110-ct-segmentation-quantize-nncf-with-output_files/index.html b/docs/notebooks/110-ct-segmentation-quantize-nncf-with-output_files/index.html index 887445ae45a..ca5e1d853de 100644 --- a/docs/notebooks/110-ct-segmentation-quantize-nncf-with-output_files/index.html +++ b/docs/notebooks/110-ct-segmentation-quantize-nncf-with-output_files/index.html @@ -1,9 +1,9 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/110-ct-segmentation-quantize-nncf-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/110-ct-segmentation-quantize-nncf-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/110-ct-segmentation-quantize-nncf-with-output_files/


../
-110-ct-segmentation-quantize-nncf-with-output_1..> 12-Jul-2023 00:11              158997
-110-ct-segmentation-quantize-nncf-with-output_3..> 12-Jul-2023 00:11              385435
-110-ct-segmentation-quantize-nncf-with-output_4..> 12-Jul-2023 00:11               73812
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/110-ct-segmentation-quantize-nncf-with-output_files/


../
+110-ct-segmentation-quantize-nncf-with-output_1..> 16-Aug-2023 01:31              158997
+110-ct-segmentation-quantize-nncf-with-output_3..> 16-Aug-2023 01:31              383352
+110-ct-segmentation-quantize-nncf-with-output_4..> 16-Aug-2023 01:31               73812
 

diff --git a/docs/notebooks/111-yolov5-quantization-migration-with-output.rst b/docs/notebooks/111-yolov5-quantization-migration-with-output.rst index 06c34c21baf..88ef1edc354 100644 --- a/docs/notebooks/111-yolov5-quantization-migration-with-output.rst +++ b/docs/notebooks/111-yolov5-quantization-migration-with-output.rst @@ -1,16 +1,15 @@ Migrate quantization from POT API to NNCF API ============================================= +.. _top: + This tutorial demonstrates how to migrate quantization pipeline written -using the OpenVINO `Post-Training Optimization Tool -(POT) `__ to -`NNCF Post-Training Quantization -API `__. -This tutorial is based on `Ultralytics -Yolov5 `__ model and additionally +using the OpenVINO `Post-Training Optimization Tool (POT) `__ to +`NNCF Post-Training Quantization API `__. +This tutorial is based on `Ultralytics YOLOv5 `__ model and additionally it compares model accuracy between the FP32 precision and quantized INT8 precision models and runs a demo of model inference based on sample code -from `Ultralytics Yolov5 `__ with +from `Ultralytics YOLOv5 `__ with the OpenVINO backend. The tutorial consists from the following parts: @@ -21,17 +20,48 @@ The tutorial consists from the following parts: 4. Perform model optimization. 5. Compare accuracy FP32 and INT8 models 6. Run model inference demo -7. Compare performance FP32 and INt8 models +7. Compare performance FP32 and INT8 models -Preparation ------------ -Download the YOLOv5 model -~~~~~~~~~~~~~~~~~~~~~~~~~ +**Table of contents**: + +- `Preparation <#preparation>`__ + + - `Download the YOLOv5 model <#download-the-yolov5-model>`__ + - `Conversion of the YOLOv5 model to OpenVINO <#conversion-of-the-yolov5-model-to-openvino>`__ + - `Imports <#imports>`__ + +- `Prepare dataset for quantization <#prepare-dataset-for-quantization>`__ + + - `Create YOLOv5 DataLoader class for POT <#create-yolov5-dataloader-class-for-pot>`__ + - `Create NNCF Dataset <#create-nncf-dataset>`__ + +- `Configure quantization pipeline <#configure-quantization-pipeline>`__ + + - `Prepare config and pipeline for POT <#prepare-config-and-pipeline-for-pot>`__ + - `Prepare configuration parameters for NNCF <#prepare-configuration-parameters-for-nncf>`__ + +- `Perform model optimization <#perform-model-optimization>`__ + + - `Run quantization using POT <#run-quantization-using-pot>`__ + - `Run quantization using NNCF <#run-quantization-using-nncf>`__ + +- `Compare accuracy FP32 and INT8 models <#compare-accuracy-fp32-and-int8-models>`__ +- `Inference Demo Performance Comparison <#inference-demo-performance-comparison>`__ +- `Benchmark <#benchmark>`__ +- `References <#references>`__ + +Preparation `⇑ <#top>`__ +############################################################################################################################### + + +Download the YOLOv5 model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 - !pip install -q 'openvino-dev>=2023.0.0' 'nncf>=2.5.0' + !pip install -q "openvino-dev>=2023.0.0" "nncf>=2.5.0" !pip install -q psutil "seaborn>=0.11.0" matplotlib numpy onnx .. code:: ipython3 @@ -63,24 +93,26 @@ Download the YOLOv5 model ``git clone https://github.com/ultralytics/yolov5.git -b v7.0`` -Conversion of the YOLOv5 model to OpenVINO -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Conversion of the YOLOv5 model to OpenVINO `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -There are three variables provided for easy run through all the notebook -cells. +There are three variables provided for easy run through all the notebook cells. -* ``IMAGE_SIZE`` - the image size for model input. -* ``MODEL_NAME`` - the model you want to use. It can be either yolov5s, yolov5m or yolov5l and so on. -* ``MODEL_PATH`` - to the path of the model directory in the YOLOv5 repository. +- ``IMAGE_SIZE`` - the image size for model input. +- ``MODEL_NAME`` - the model you want to use. It can be either yolov5s, + yolov5m or yolov5l and so on. +- ``MODEL_PATH`` - to the path of the model directory in the YOLOv5 + repository. -YoloV5 ``export.py`` scripts support multiple model formats for +YOLOv5 ``export.py`` scripts support multiple model formats for conversion. ONNX is also represented among supported formats. We need to specify ``--include ONNX`` parameter for exporting. As the result, directory with the ``{MODEL_NAME}`` name will be created with the -following content: +following content: -* ``{MODEL_NAME}.pt`` - the downloaded pre-trained weight. -* ``{MODEL_NAME}.onnx`` - the Open Neural Network Exchange (ONNX) is an open format, built to represent machine learning models. +- ``{MODEL_NAME}.pt`` - the downloaded pre-trained weight. +- ``{MODEL_NAME}.onnx`` - the Open Neural Network Exchange (ONNX) is an + open format, built to represent machine learning models. .. code:: ipython3 @@ -111,7 +143,7 @@ following content: YOLOv5 🚀 v7.0-0-g915bbf2 Python-3.8.10 torch-1.13.1+cpu CPU Downloading https://github.com/ultralytics/yolov5/releases/download/v7.0/yolov5m.pt to yolov5m/yolov5m.pt... - 100%|██████████████████████████████████████| 40.8M/40.8M [00:09<00:00, 4.52MB/s] + 100%|██████████████████████████████████████| 40.8M/40.8M [00:10<00:00, 4.02MB/s] Fusing layers... YOLOv5m summary: 290 layers, 21172173 parameters, 0 gradients @@ -119,10 +151,10 @@ following content: PyTorch: starting from yolov5m/yolov5m.pt with output shape (1, 25200, 85) (40.8 MB) ONNX: starting export with onnx 1.14.0... - ONNX: export success ✅ 1.2s, saved as yolov5m/yolov5m.onnx (81.2 MB) + ONNX: export success ✅ 1.3s, saved as yolov5m/yolov5m.onnx (81.2 MB) - Export complete (12.2s) - Results saved to /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/yolov5/yolov5m + Export complete (13.7s) + Results saved to /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/yolov5/yolov5m Detect: python detect.py --weights yolov5m/yolov5m.onnx Validate: python val.py --weights yolov5m/yolov5m.onnx PyTorch Hub: model = torch.hub.load('ultralytics/yolov5', 'custom', 'yolov5m/yolov5m.onnx') @@ -130,17 +162,17 @@ following content: Convert the ONNX model to OpenVINO Intermediate Representation (IR) -model generated by `Model -Optimizer `__. -We will use `Model Optimizer Python -API `__ -``openvino.tools.mo.convert_model`` function to convert ONNX model to -OpenVINO Model, then it can be seralized using -``openvino.runtime.serialize``\ As the result, directory with the -``{MODEL_DIR}`` name will be created with the following content: - -* ``{MODEL_NAME}_fp32.xml``, ``{MODEL_NAME}_fp32.bin`` - OpenVINO Intermediate Representation (IR) model format with FP32 precision generated by Model Optimizer. -* ``{MODEL_NAME}_fp16.xml``, ``{MODEL_NAME}_fp16.bin`` - OpenVINO Intermediate Representation (IR) model format with FP32 precision generated by Model Optimizer. +model generated by `model conversion API `__. +We will use the ``openvino.tools.mo.convert_model`` function of model +conversion Python API to convert ONNX model to OpenVINO Model, then it +can be serialized using ``openvino.runtime.serialize``. As the result, +directory with the ``{MODEL_DIR}`` name will be created with the +following content: \* ``{MODEL_NAME}_fp32.xml``, +``{MODEL_NAME}_fp32.bin`` - OpenVINO Intermediate Representation (IR) +model format with FP32 precision generated by Model Optimizer. \* +``{MODEL_NAME}_fp16.xml``, ``{MODEL_NAME}_fp16.bin`` - OpenVINO +Intermediate Representation (IR) model format with FP32 precision +generated by Model Optimizer. .. code:: ipython3 @@ -170,8 +202,9 @@ OpenVINO Model, then it can be seralized using Export ONNX to OpenVINO FP16 IR to: yolov5/yolov5m/FP16_openvino_model/yolov5m_fp16.xml -Imports -~~~~~~~ +Imports `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -180,12 +213,13 @@ Imports from yolov5.utils.dataloaders import create_dataloader from yolov5.utils.general import check_dataset -Prepare dataset for quantization --------------------------------- +Prepare dataset for quantization `⇑ <#top>`__ +############################################################################################################################### + Before starting quantization, we should prepare dataset, which will be used for quantization. Ultralytics YOLOv5 provides data loader for -iteration overdataset during training and validation. Let’s create it +iteration over dataset during training and validation. Let’s create it first. .. code:: ipython3 @@ -229,12 +263,13 @@ first. .. parsed-literal:: Unzipping datasets/coco128.zip... - Scanning /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/datasets/coco128/labels/train2017... 126 images, 2 backgrounds, 0 corrupt: 100%|██████████| 128/128 00:00 - New cache created: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/datasets/coco128/labels/train2017.cache + Scanning /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/datasets/coco128/labels/train2017... 126 images, 2 backgrounds, 0 corrupt: 100%|██████████| 128/128 00:00 + New cache created: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/datasets/coco128/labels/train2017.cache -Create YOLOv5 DataLoader class for POT -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Create YOLOv5 DataLoader class for POT `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Create a class for loading the YOLOv5 dataset and annotation which inherits from POT API class DataLoader. @@ -245,15 +280,15 @@ index. Any implementation should override the following methods: - The ``__len__()``, returns the size of the dataset. - The ``__getitem__()``, provides access to the data by index in range - of 0 to len(self). It can also encapsulate the logic of + of 0 to ``len(self)``. It can also encapsulate the logic of model-specific pre-processing. This method should return data in the (data, annotation) format, in which: - The ``data`` is the input that is passed to the model at inference so that it should be properly preprocessed. It can be either the - numpy.array object or a dictionary, where the key is the name of - the model input and value is numpy.array which corresponds to this - input. + ``numpy.array`` object or a dictionary, where the key is the name + of the model input and value is ``numpy.array`` which corresponds + to this input. - The ``annotation`` is not used by the Default Quantization method. Therefore, this object can be None in this case. @@ -304,7 +339,7 @@ index. Any implementation should override the following methods: .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/offline_transformations/__init__.py:10: FutureWarning: The module is private and following namespace `offline_transformations` will be removed in the future. + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/offline_transformations/__init__.py:10: FutureWarning: The module is private and following namespace `offline_transformations` will be removed in the future. warnings.warn( @@ -322,15 +357,16 @@ index. Any implementation should override the following methods: Nevergrad package could not be imported. If you are planning to use any hyperparameter optimization algo, consider installing it using pip. This implies advanced usage of the tool. Note that nevergrad is compatible only with Python 3.7+ -Create NNCF Dataset -~~~~~~~~~~~~~~~~~~~ +Create NNCF Dataset `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + For preparing quantization dataset for NNCF, we should wrap framework-specific data source into ``nncf.Dataset`` instance. -Additionaly, to transform data into model expected format we can define +Additionally, to transform data into model expected format we can define transformation function, which accept data item for single dataset -iteration and transform it for feeding into model (e.g. in simpliest -case, if data item contains input tensor and anntation, we should +iteration and transform it for feeding into model (e.g. in simplest +case, if data item contains input tensor and annotation, we should extract only input data from it and convert it into model expected format). @@ -361,13 +397,15 @@ format). INFO:nncf:NNCF initialized successfully. Supported frameworks detected: torch, tensorflow, onnx, openvino -Configure quantization pipeline -------------------------------- +Configure quantization pipeline `⇑ <#top>`__ +############################################################################################################################### + Next, we should define quantization algorithm parameters. -Prepare config and pipeline for POT -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Prepare config and pipeline for POT `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + in POT, all quantization parameters should be defined using configuration dictionary. Config consists of 3 sections: ``algorithms`` @@ -415,12 +453,13 @@ pipeline using ``create_pipeline`` function. # Step 5: Create a pipeline of compression algorithms. pipeline = create_pipeline(algorithms_config, engine) -Prapare configuration parameters for NNCF -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Prepare configuration parameters for NNCF `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Post-training quantization pipeline in NNCF represented by -``nncf.quantize`` function for DefaultQuantization Algorithm and -``nncf.quantize_with_accuracy_control`` for AccuracyAwareQuantization. +``nncf.quantize`` function for Default Quantization Algorithm and +``nncf.quantize_with_accuracy_control`` for Accuracy Aware Quantization. Quantization parameters ``preset``, ``model_type``, ``subset_size``, ``fast_bias_correction``, ``ignored_scope`` are arguments of function. More details about supported parameters and formats can be found in NNCF @@ -435,11 +474,13 @@ in our case ``openvino.runtime.Model`` instance created using subset_size = 300 preset = nncf.QuantizationPreset.MIXED -Perform model optimization --------------------------- +Perform model optimization `⇑ <#top>`__ +############################################################################################################################### + + +Run quantization using POT `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Run quantization using POT -~~~~~~~~~~~~~~~~~~~~~~~~~~ To start model quantization using POT API, we should call ``pipeline.run(pot_model)`` method. As the result, we got quantized @@ -459,13 +500,14 @@ size of final .bin file. save_model(compressed_model, optimized_save_dir, model_config["model_name"] + "_int8") pot_int8_path = f"{optimized_save_dir}/{MODEL_NAME}_int8.xml" -Run quantization using NNCF -~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Run quantization using NNCF `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + To run NNCF quantization, we should call ``nncf.quantize`` function. As the result, the function returns quantized model in the same format like input model, so it means that quantized model ready to be compiled on -device for inferece and can be saved on disk using +device for inference and can be saved on disk using ``openvino.runtime.serialize``. .. code:: ipython3 @@ -483,16 +525,17 @@ device for inferece and can be saved on disk using .. parsed-literal:: - Statistics collection: 43%|████▎ | 128/300 [00:30<00:41, 4.17it/s] - Biases correction: 100%|██████████| 82/82 [00:10<00:00, 7.86it/s] + Statistics collection: 43%|████▎ | 128/300 [00:30<00:40, 4.20it/s] + Biases correction: 100%|██████████| 82/82 [00:10<00:00, 7.71it/s] -Compare accuracy FP32 and INT8 models -------------------------------------- +Compare accuracy FP32 and INT8 models `⇑ <#top>`__ +############################################################################################################################### + For getting accuracy results, we will use ``yolov5.val.run`` function which already supports OpenVINO backend. For making int8 model is -compatible with Ultralytics provided validation pipeline, we alse should +compatible with Ultralytics provided validation pipeline, we also should provide metadata with information about supported class names in the same directory, where model located. @@ -551,10 +594,10 @@ same directory, where model located. .. parsed-literal:: Forcing --batch-size 1 square inference (1,3,640,640) for non-PyTorch models - val: Scanning /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/datasets/coco128/labels/train2017.cache... 126 images, 2 backgrounds, 0 corrupt: 100%|██████████| 128/128 00:00 + val: Scanning /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/datasets/coco128/labels/train2017.cache... 126 images, 2 backgrounds, 0 corrupt: 100%|██████████| 128/128 00:00 Class Images Instances P R mAP50 mAP50-95: 100%|██████████| 128/128 00:05 all 128 929 0.726 0.687 0.769 0.554 - Speed: 0.2ms pre-process, 35.3ms inference, 3.2ms NMS per image at shape (1, 3, 640, 640) + Speed: 0.2ms pre-process, 35.3ms inference, 3.0ms NMS per image at shape (1, 3, 640, 640) Results saved to yolov5/runs/val/exp @@ -598,10 +641,10 @@ same directory, where model located. .. parsed-literal:: Forcing --batch-size 1 square inference (1,3,640,640) for non-PyTorch models - val: Scanning /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/datasets/coco128/labels/train2017.cache... 126 images, 2 backgrounds, 0 corrupt: 100%|██████████| 128/128 00:00 + val: Scanning /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/datasets/coco128/labels/train2017.cache... 126 images, 2 backgrounds, 0 corrupt: 100%|██████████| 128/128 00:00 Class Images Instances P R mAP50 mAP50-95: 100%|██████████| 128/128 00:03 all 128 929 0.761 0.677 0.773 0.548 - Speed: 0.2ms pre-process, 17.3ms inference, 3.3ms NMS per image at shape (1, 3, 640, 640) + Speed: 0.2ms pre-process, 17.1ms inference, 3.3ms NMS per image at shape (1, 3, 640, 640) Results saved to yolov5/runs/val/exp2 @@ -645,10 +688,10 @@ same directory, where model located. .. parsed-literal:: Forcing --batch-size 1 square inference (1,3,640,640) for non-PyTorch models - val: Scanning /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/datasets/coco128/labels/train2017.cache... 126 images, 2 backgrounds, 0 corrupt: 100%|██████████| 128/128 00:00 + val: Scanning /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/datasets/coco128/labels/train2017.cache... 126 images, 2 backgrounds, 0 corrupt: 100%|██████████| 128/128 00:00 Class Images Instances P R mAP50 mAP50-95: 100%|██████████| 128/128 00:03 all 128 929 0.742 0.684 0.766 0.546 - Speed: 0.2ms pre-process, 17.1ms inference, 3.3ms NMS per image at shape (1, 3, 640, 640) + Speed: 0.2ms pre-process, 17.0ms inference, 3.2ms NMS per image at shape (1, 3, 640, 640) Results saved to yolov5/runs/val/exp3 @@ -708,14 +751,15 @@ model. -.. image:: 111-yolov5-quantization-migration-with-output_files/111-yolov5-quantization-migration-with-output_33_0.png +.. image:: 111-yolov5-quantization-migration-with-output_files/111-yolov5-quantization-migration-with-output_34_0.png -Inference Demo Performance Comparison -------------------------------------- +Inference Demo Performance Comparison `⇑ <#top>`__ +############################################################################################################################### -This part shows how to use the Ultralytics model detection code -`“detect.py” `__ + +This part shows how to use the Ultralytics model detection code +`detect.py `__ to run synchronous inference, using the OpenVINO Python API on two images. @@ -744,9 +788,9 @@ images. 'YOLOv5 🚀 v7.0-0-g915bbf2 Python-3.8.10 torch-1.13.1+cpu CPU', '', 'Loading yolov5m/FP32_openvino_model for OpenVINO inference...', - 'image 1/2 /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/yolov5/data/images/bus.jpg: 640x640 4 persons, 1 bus, 56.6ms', - 'image 2/2 /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/yolov5/data/images/zidane.jpg: 640x640 3 persons, 2 ties, 43.1ms', - 'Speed: 1.4ms pre-process, 49.8ms inference, 1.2ms NMS per image at shape (1, 3, 640, 640)', + 'image 1/2 /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/yolov5/data/images/bus.jpg: 640x640 4 persons, 1 bus, 57.0ms', + 'image 2/2 /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/yolov5/data/images/zidane.jpg: 640x640 3 persons, 2 ties, 40.6ms', + 'Speed: 1.4ms pre-process, 48.8ms inference, 1.2ms NMS per image at shape (1, 3, 640, 640)', 'Results saved to \x1b[1mruns/detect/exp\x1b[0m'] @@ -771,9 +815,9 @@ images. 'YOLOv5 🚀 v7.0-0-g915bbf2 Python-3.8.10 torch-1.13.1+cpu CPU', '', 'Loading yolov5m/POT_INT8_openvino_model for OpenVINO inference...', - 'image 1/2 /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/yolov5/data/images/bus.jpg: 640x640 4 persons, 1 bus, 38.2ms', - 'image 2/2 /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/yolov5/data/images/zidane.jpg: 640x640 3 persons, 1 tie, 33.4ms', - 'Speed: 1.6ms pre-process, 35.8ms inference, 1.4ms NMS per image at shape (1, 3, 640, 640)', + 'image 1/2 /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/yolov5/data/images/bus.jpg: 640x640 4 persons, 1 bus, 36.6ms', + 'image 2/2 /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/yolov5/data/images/zidane.jpg: 640x640 3 persons, 1 tie, 33.4ms', + 'Speed: 1.5ms pre-process, 35.0ms inference, 1.4ms NMS per image at shape (1, 3, 640, 640)', 'Results saved to \x1b[1mruns/detect/exp2\x1b[0m'] @@ -798,9 +842,9 @@ images. 'YOLOv5 🚀 v7.0-0-g915bbf2 Python-3.8.10 torch-1.13.1+cpu CPU', '', 'Loading yolov5m/NNCF_INT8_openvino_model for OpenVINO inference...', - 'image 1/2 /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/yolov5/data/images/bus.jpg: 640x640 4 persons, 1 bus, 37.5ms', - 'image 2/2 /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/yolov5/data/images/zidane.jpg: 640x640 3 persons, 2 ties, 32.4ms', - 'Speed: 1.5ms pre-process, 35.0ms inference, 1.3ms NMS per image at shape (1, 3, 640, 640)', + 'image 1/2 /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/yolov5/data/images/bus.jpg: 640x640 4 persons, 1 bus, 35.9ms', + 'image 2/2 /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/111-yolov5-quantization-migration/yolov5/data/images/zidane.jpg: 640x640 3 persons, 2 ties, 31.5ms', + 'Speed: 1.6ms pre-process, 33.7ms inference, 1.4ms NMS per image at shape (1, 3, 640, 640)', 'Results saved to \x1b[1mruns/detect/exp3\x1b[0m'] @@ -828,11 +872,12 @@ images. -.. image:: 111-yolov5-quantization-migration-with-output_files/111-yolov5-quantization-migration-with-output_39_0.png +.. image:: 111-yolov5-quantization-migration-with-output_files/111-yolov5-quantization-migration-with-output_40_0.png -Benchmark ---------- +Benchmark `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -864,7 +909,7 @@ Benchmark [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.THROUGHPUT. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 34.56 ms + [ INFO ] Read model took 31.30 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] images (node: images) : f32 / [...] / [1,3,640,640] @@ -878,7 +923,7 @@ Benchmark [ INFO ] Model outputs: [ INFO ] output0 (node: output0) : f32 / [...] / [1,25200,85] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 382.88 ms + [ INFO ] Compile model took 360.18 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: [ INFO ] NETWORK_NAME: torch_jit @@ -900,17 +945,17 @@ Benchmark [ INFO ] Fill input 'images' with random values [Step 10/11] Measuring performance (Start inference asynchronously, 6 inference requests, limits: 15000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 106.76 ms + [ INFO ] First inference took 102.48 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 450 iterations - [ INFO ] Duration: 15298.04 ms + [ INFO ] Count: 456 iterations + [ INFO ] Duration: 15352.58 ms [ INFO ] Latency: - [ INFO ] Median: 204.03 ms - [ INFO ] Average: 202.87 ms - [ INFO ] Min: 140.18 ms - [ INFO ] Max: 217.69 ms - [ INFO ] Throughput: 29.42 FPS + [ INFO ] Median: 202.33 ms + [ INFO ] Average: 201.47 ms + [ INFO ] Min: 138.12 ms + [ INFO ] Max: 216.53 ms + [ INFO ] Throughput: 29.70 FPS .. code:: ipython3 @@ -941,7 +986,7 @@ Benchmark [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.THROUGHPUT. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 39.27 ms + [ INFO ] Read model took 44.00 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] images (node: images) : f32 / [...] / [1,3,640,640] @@ -955,7 +1000,7 @@ Benchmark [ INFO ] Model outputs: [ INFO ] output0 (node: output0) : f32 / [...] / [1,25200,85] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 409.09 ms + [ INFO ] Compile model took 387.13 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: [ INFO ] NETWORK_NAME: torch_jit @@ -977,17 +1022,17 @@ Benchmark [ INFO ] Fill input 'images' with random values [Step 10/11] Measuring performance (Start inference asynchronously, 6 inference requests, limits: 15000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 103.11 ms + [ INFO ] First inference took 101.31 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] [ INFO ] Count: 456 iterations - [ INFO ] Duration: 15294.88 ms + [ INFO ] Duration: 15246.15 ms [ INFO ] Latency: - [ INFO ] Median: 201.67 ms - [ INFO ] Average: 200.70 ms - [ INFO ] Min: 124.86 ms - [ INFO ] Max: 218.07 ms - [ INFO ] Throughput: 29.81 FPS + [ INFO ] Median: 200.30 ms + [ INFO ] Average: 199.86 ms + [ INFO ] Min: 96.50 ms + [ INFO ] Max: 219.99 ms + [ INFO ] Throughput: 29.91 FPS .. code:: ipython3 @@ -1018,7 +1063,7 @@ Benchmark [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.THROUGHPUT. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 47.44 ms + [ INFO ] Read model took 45.57 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] images (node: images) : f32 / [...] / [1,3,640,640] @@ -1032,7 +1077,7 @@ Benchmark [ INFO ] Model outputs: [ INFO ] output0 (node: output0) : f32 / [...] / [1,25200,85] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 705.56 ms + [ INFO ] Compile model took 702.15 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: [ INFO ] NETWORK_NAME: torch_jit @@ -1054,17 +1099,17 @@ Benchmark [ INFO ] Fill input 'images' with random values [Step 10/11] Measuring performance (Start inference asynchronously, 6 inference requests, limits: 15000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 50.77 ms + [ INFO ] First inference took 48.39 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 1422 iterations - [ INFO ] Duration: 15093.38 ms + [ INFO ] Count: 1410 iterations + [ INFO ] Duration: 15055.06 ms [ INFO ] Latency: - [ INFO ] Median: 63.57 ms - [ INFO ] Average: 63.51 ms - [ INFO ] Min: 44.75 ms - [ INFO ] Max: 86.46 ms - [ INFO ] Throughput: 94.21 FPS + [ INFO ] Median: 64.05 ms + [ INFO ] Average: 63.87 ms + [ INFO ] Min: 46.43 ms + [ INFO ] Max: 83.05 ms + [ INFO ] Throughput: 93.66 FPS .. code:: ipython3 @@ -1095,7 +1140,7 @@ Benchmark [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.THROUGHPUT. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 53.05 ms + [ INFO ] Read model took 52.31 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] images (node: images) : f32 / [...] / [1,3,640,640] @@ -1109,7 +1154,7 @@ Benchmark [ INFO ] Model outputs: [ INFO ] output0 (node: output0) : f32 / [...] / [1,25200,85] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 714.97 ms + [ INFO ] Compile model took 710.35 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: [ INFO ] NETWORK_NAME: torch_jit @@ -1131,26 +1176,27 @@ Benchmark [ INFO ] Fill input 'images' with random values [Step 10/11] Measuring performance (Start inference asynchronously, 6 inference requests, limits: 15000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 53.03 ms + [ INFO ] First inference took 51.01 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 1422 iterations - [ INFO ] Duration: 15073.87 ms + [ INFO ] Count: 1416 iterations + [ INFO ] Duration: 15113.02 ms [ INFO ] Latency: - [ INFO ] Median: 63.63 ms - [ INFO ] Average: 63.46 ms - [ INFO ] Min: 53.07 ms - [ INFO ] Max: 85.38 ms - [ INFO ] Throughput: 94.34 FPS + [ INFO ] Median: 63.94 ms + [ INFO ] Average: 63.87 ms + [ INFO ] Min: 45.06 ms + [ INFO ] Max: 87.95 ms + [ INFO ] Throughput: 93.69 FPS -References ----------- +References `⇑ <#top>`__ +############################################################################################################################### + - `Ultralytics YOLOv5 `__ - `OpenVINO Post-training Optimization Tool `__ - `NNCF Post-training - quantization `__ -- `Model - Optimizer `__ + quantization `__ +- `Model Conversion + API `__ diff --git a/docs/notebooks/111-yolov5-quantization-migration-with-output_files/111-yolov5-quantization-migration-with-output_33_0.png b/docs/notebooks/111-yolov5-quantization-migration-with-output_files/111-yolov5-quantization-migration-with-output_34_0.png similarity index 100% rename from docs/notebooks/111-yolov5-quantization-migration-with-output_files/111-yolov5-quantization-migration-with-output_33_0.png rename to docs/notebooks/111-yolov5-quantization-migration-with-output_files/111-yolov5-quantization-migration-with-output_34_0.png diff --git a/docs/notebooks/111-yolov5-quantization-migration-with-output_files/111-yolov5-quantization-migration-with-output_39_0.png b/docs/notebooks/111-yolov5-quantization-migration-with-output_files/111-yolov5-quantization-migration-with-output_40_0.png similarity index 100% rename from docs/notebooks/111-yolov5-quantization-migration-with-output_files/111-yolov5-quantization-migration-with-output_39_0.png rename to docs/notebooks/111-yolov5-quantization-migration-with-output_files/111-yolov5-quantization-migration-with-output_40_0.png diff --git a/docs/notebooks/111-yolov5-quantization-migration-with-output_files/index.html b/docs/notebooks/111-yolov5-quantization-migration-with-output_files/index.html index fe5e236eb3d..d42c2759d7f 100644 --- a/docs/notebooks/111-yolov5-quantization-migration-with-output_files/index.html +++ b/docs/notebooks/111-yolov5-quantization-migration-with-output_files/index.html @@ -1,8 +1,8 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/111-yolov5-quantization-migration-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/111-yolov5-quantization-migration-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/111-yolov5-quantization-migration-with-output_files/


../
-111-yolov5-quantization-migration-with-output_3..> 12-Jul-2023 00:11               33667
-111-yolov5-quantization-migration-with-output_3..> 12-Jul-2023 00:11              770524
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/111-yolov5-quantization-migration-with-output_files/


../
+111-yolov5-quantization-migration-with-output_3..> 16-Aug-2023 01:31               33667
+111-yolov5-quantization-migration-with-output_4..> 16-Aug-2023 01:31              770524
 

diff --git a/docs/notebooks/112-pytorch-post-training-quantization-nncf-with-output.rst b/docs/notebooks/112-pytorch-post-training-quantization-nncf-with-output.rst index ddc6b8bc22f..b4e5caafbb3 100644 --- a/docs/notebooks/112-pytorch-post-training-quantization-nncf-with-output.rst +++ b/docs/notebooks/112-pytorch-post-training-quantization-nncf-with-output.rst @@ -1,6 +1,8 @@ Post-Training Quantization of PyTorch models with NNCF ====================================================== +.. _top: + The goal of this tutorial is to demonstrate how to use the NNCF (Neural Network Compression Framework) 8-bit quantization in post-training mode (without the fine-tuning pipeline) to optimize a PyTorch model for the @@ -20,10 +22,31 @@ quantization, not demanding the fine-tuning of the model. **NOTE**: This notebook requires that a C++ compiler is accessible on the default binary search path of the OS you are running the - notebook. + notebook. + + +**Table of contents**: + +- `Preparations <#preparations>`__ + + - `Imports <#imports>`__ + - `Settings <#settings>`__ + - `Download and Prepare Tiny ImageNet dataset <#download-and-prepare-tiny-imagenet-dataset>`__ + - `Helpers classes and functions <#helpers-classes-and-functions>`__ + - `Validation function <#validation-function>`__ + - `Create and load original uncompressed model <#create-and-load-original-uncompressed-model>`__ + - `Create train and validation DataLoaders <#create-train-and-validation-dataloaders>`__ + +- `Model quantization and benchmarking <#model-quantization-and-benchmarking>`__ + + - `I. Evaluate the loaded model <#i-evaluate-the-loaded-model>`__ + - `II. Create and initialize quantization <#ii-create-and-initialize-quantization>`__ + - `III. Convert the models to OpenVINO Intermediate Representation (OpenVINO IR) <#iii-convert-the-models-to-openvino-intermediate-representation-openvino-ir>`__ + - `IV. Compare performance of INT8 model and FP32 model in OpenVINO <#iv-compare-performance-of-int8-model-and-fp32-model-in-openvino>`__ + +Preparations `⇑ <#top>`__ +############################################################################################################################### -Preparations ------------- .. code:: ipython3 @@ -62,8 +85,9 @@ Preparations os.environ["LIB"] = os.pathsep.join(b.library_dirs) print(f"Added {vs_dir} to PATH") -Imports -~~~~~~~ +Imports `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -88,10 +112,10 @@ Imports .. parsed-literal:: - 2023-07-11 22:44:00.778229: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. - 2023-07-11 22:44:00.812197: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + 2023-08-15 22:47:54.862445: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-08-15 22:47:54.896717: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. - 2023-07-11 22:44:01.354577: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + 2023-08-15 22:47:55.440534: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT .. parsed-literal:: @@ -99,8 +123,9 @@ Imports INFO:nncf:NNCF initialized successfully. Supported frameworks detected: torch, tensorflow, onnx, openvino -Settings -~~~~~~~~ +Settings `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -142,12 +167,13 @@ Settings .. parsed-literal:: - PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/112-pytorch-post-training-quantization-nncf/model/resnet50_fp32.pth') + PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/112-pytorch-post-training-quantization-nncf/model/resnet50_fp32.pth') -Download and Prepare Tiny ImageNet dataset -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Download and Prepare Tiny ImageNet dataset `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + - 100k images of shape 3x64x64, - 200 different classes: snake, spider, cat, truck, grasshopper, gull, @@ -206,11 +232,10 @@ Download and Prepare Tiny ImageNet dataset Successfully downloaded and extracted dataset to: output -Helpers classes and functions -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Helpers classes and functions `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -The code below will help to count accuracy and visualize validation -process. +The code below will help to count accuracy and visualize validation process. .. code:: ipython3 @@ -272,8 +297,9 @@ process. return res -Validation function -~~~~~~~~~~~~~~~~~~~ +Validation function `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -325,10 +351,11 @@ Validation function ) return top1.avg -Create and load original uncompressed model -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Create and load original uncompressed model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -ResNet-50 from the `torchivision + +ResNet-50 from the ```torchivision`` repository `__ is pre-trained on ImageNet with more prediction classes than Tiny ImageNet, so the model is adjusted by swapping the last FC layer to one with fewer output @@ -353,8 +380,9 @@ values. model = create_model(MODEL_DIR / fp32_checkpoint_filename) -Create train and validation dataloaders -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Create train and validation DataLoaders `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -403,15 +431,16 @@ Create train and validation dataloaders train_loader, val_loader = create_dataloaders() -Model quantization and benchmarking ------------------------------------ +Model quantization and benchmarking `⇑ <#top>`__ +############################################################################################################################### -With the validation pipeline, model files, and data-loading procedures -for model calibration now prepared, it’s time to proceed with the actual -post-training quantization using NNCF. +With the validation pipeline, model files, and data-loading procedures for model calibration +now prepared, it’s time to proceed with the actual post-training +quantization using NNCF. + +I. Evaluate the loaded model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -I. Evaluate the loaded model -~~~~~~~~~~~~~~~~~~~~~~~~~~~~ .. code:: ipython3 @@ -421,28 +450,28 @@ I. Evaluate the loaded model .. parsed-literal:: - Test: [ 0/79] Time 0.253 (0.253) Acc@1 81.25 (81.25) Acc@5 92.19 (92.19) - Test: [10/79] Time 0.238 (0.240) Acc@1 56.25 (66.97) Acc@5 86.72 (87.50) - Test: [20/79] Time 0.235 (0.239) Acc@1 67.97 (64.29) Acc@5 85.16 (87.35) - Test: [30/79] Time 0.234 (0.238) Acc@1 53.12 (62.37) Acc@5 77.34 (85.33) - Test: [40/79] Time 0.235 (0.238) Acc@1 67.19 (60.86) Acc@5 90.62 (84.51) - Test: [50/79] Time 0.232 (0.240) Acc@1 60.16 (60.80) Acc@5 88.28 (84.42) - Test: [60/79] Time 0.251 (0.240) Acc@1 66.41 (60.46) Acc@5 86.72 (83.79) - Test: [70/79] Time 0.237 (0.240) Acc@1 52.34 (60.21) Acc@5 80.47 (83.33) - * Acc@1 60.740 Acc@5 83.960 Total time: 18.698 + Test: [ 0/79] Time 0.240 (0.240) Acc@1 81.25 (81.25) Acc@5 92.19 (92.19) + Test: [10/79] Time 0.234 (0.227) Acc@1 56.25 (66.97) Acc@5 86.72 (87.50) + Test: [20/79] Time 0.220 (0.225) Acc@1 67.97 (64.29) Acc@5 85.16 (87.35) + Test: [30/79] Time 0.219 (0.223) Acc@1 53.12 (62.37) Acc@5 77.34 (85.33) + Test: [40/79] Time 0.225 (0.222) Acc@1 67.19 (60.86) Acc@5 90.62 (84.51) + Test: [50/79] Time 0.220 (0.222) Acc@1 60.16 (60.80) Acc@5 88.28 (84.42) + Test: [60/79] Time 0.219 (0.222) Acc@1 66.41 (60.46) Acc@5 86.72 (83.79) + Test: [70/79] Time 0.219 (0.222) Acc@1 52.34 (60.21) Acc@5 80.47 (83.33) + * Acc@1 60.740 Acc@5 83.960 Total time: 17.387 Test accuracy of FP32 model: 60.740 -II. Create and initialize quantization -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +II. Create and initialize quantization `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -NNCF enables post-training quantization by adding the quantization -layers into the model graph and then using a subset of the training -dataset to initialize the parameters of these additional quantization -layers. The framework is designed so that modifications to your original -training code are minor. Quantization is the simplest scenario and -requires a few modifications. For more information about NNCF Post -Training Quantization (PTQ) API, refer to the `Basic Quantization Flow +NNCF enables post-training quantization by adding the quantization layers into the +model graph and then using a subset of the training dataset to +initialize the parameters of these additional quantization layers. The +framework is designed so that modifications to your original training +code are minor. Quantization is the simplest scenario and requires a few +modifications. For more information about NNCF Post Training +Quantization (PTQ) API, refer to the `Basic Quantization Flow Guide `__. 1. Create a transformation function that accepts a sample from the @@ -498,16 +527,16 @@ Guide `__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Use Model Optimizer Python API to convert the Pytorch models to OpenVINO -IR. The models will be saved to the ‘OUTPUT’ directory for latter -benchmarking. +To convert the Pytorch models to OpenVINO IR, use model conversion +Python API . The models will be saved to the ‘OUTPUT’ directory for +later benchmarking. -For more information about Model Optimizer, refer to the `Model -Optimizer Developer -Guide `__. +For more information about model conversion, refer to this +`page `__. -Before converting models export them to ONNX. Executing the following +Before converting models, export them to ONNX. Executing the following command may take a while. .. code:: ipython3 @@ -549,17 +577,17 @@ command may take a while. .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/nncf/torch/quantization/layers.py:338: TracerWarning: Converting a tensor to a Python number might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/nncf/torch/quantization/layers.py:338: TracerWarning: Converting a tensor to a Python number might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! return self._level_low.item() - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/nncf/torch/quantization/layers.py:346: TracerWarning: Converting a tensor to a Python number might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/nncf/torch/quantization/layers.py:346: TracerWarning: Converting a tensor to a Python number might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! return self._level_high.item() - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/nncf/torch/quantization/quantize_functions.py:140: FutureWarning: 'torch.onnx._patch_torch._graph_op' is deprecated in version 1.13 and will be removed in version 1.14. Please note 'g.op()' is to be removed from torch.Graph. Please open a GitHub issue if you need this functionality.. + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/nncf/torch/quantization/quantize_functions.py:140: FutureWarning: 'torch.onnx._patch_torch._graph_op' is deprecated in version 1.13 and will be removed in version 1.14. Please note 'g.op()' is to be removed from torch.Graph. Please open a GitHub issue if you need this functionality.. output = g.op( - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/_patch_torch.py:81: UserWarning: The shape inference of org.openvinotoolkit::FakeQuantize type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/_patch_torch.py:81: UserWarning: The shape inference of org.openvinotoolkit::FakeQuantize type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) _C._jit_pass_onnx_node_shape_type_inference( - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: The shape inference of org.openvinotoolkit::FakeQuantize type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: The shape inference of org.openvinotoolkit::FakeQuantize type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) _C._jit_pass_onnx_graph_shape_type_inference( - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: The shape inference of org.openvinotoolkit::FakeQuantize type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: The shape inference of org.openvinotoolkit::FakeQuantize type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) _C._jit_pass_onnx_graph_shape_type_inference( @@ -600,15 +628,15 @@ Evaluate the FP32 and INT8 models. .. parsed-literal:: - Test: [ 0/79] Time 0.205 (0.205) Acc@1 81.25 (81.25) Acc@5 92.19 (92.19) - Test: [10/79] Time 0.137 (0.143) Acc@1 56.25 (66.97) Acc@5 86.72 (87.50) - Test: [20/79] Time 0.145 (0.141) Acc@1 67.97 (64.29) Acc@5 85.16 (87.35) - Test: [30/79] Time 0.137 (0.140) Acc@1 53.12 (62.37) Acc@5 77.34 (85.33) - Test: [40/79] Time 0.137 (0.140) Acc@1 67.19 (60.86) Acc@5 90.62 (84.51) - Test: [50/79] Time 0.139 (0.140) Acc@1 60.16 (60.80) Acc@5 88.28 (84.42) - Test: [60/79] Time 0.135 (0.139) Acc@1 66.41 (60.46) Acc@5 86.72 (83.79) - Test: [70/79] Time 0.139 (0.139) Acc@1 52.34 (60.21) Acc@5 80.47 (83.33) - * Acc@1 60.740 Acc@5 83.960 Total time: 10.882 + Test: [ 0/79] Time 0.200 (0.200) Acc@1 81.25 (81.25) Acc@5 92.19 (92.19) + Test: [10/79] Time 0.138 (0.144) Acc@1 56.25 (66.97) Acc@5 86.72 (87.50) + Test: [20/79] Time 0.137 (0.141) Acc@1 67.97 (64.29) Acc@5 85.16 (87.35) + Test: [30/79] Time 0.136 (0.140) Acc@1 53.12 (62.37) Acc@5 77.34 (85.33) + Test: [40/79] Time 0.139 (0.140) Acc@1 67.19 (60.86) Acc@5 90.62 (84.51) + Test: [50/79] Time 0.135 (0.139) Acc@1 60.16 (60.80) Acc@5 88.28 (84.42) + Test: [60/79] Time 0.139 (0.139) Acc@1 66.41 (60.46) Acc@5 86.72 (83.79) + Test: [70/79] Time 0.138 (0.139) Acc@1 52.34 (60.21) Acc@5 80.47 (83.33) + * Acc@1 60.740 Acc@5 83.960 Total time: 10.865 Accuracy of FP32 IR model: 60.740 @@ -621,20 +649,20 @@ Evaluate the FP32 and INT8 models. .. parsed-literal:: - Test: [ 0/79] Time 0.187 (0.187) Acc@1 81.25 (81.25) Acc@5 90.62 (90.62) - Test: [10/79] Time 0.082 (0.091) Acc@1 57.81 (66.90) Acc@5 85.94 (87.78) - Test: [20/79] Time 0.080 (0.086) Acc@1 67.97 (64.06) Acc@5 83.59 (87.31) - Test: [30/79] Time 0.081 (0.084) Acc@1 52.34 (62.17) Acc@5 78.12 (85.36) - Test: [40/79] Time 0.079 (0.083) Acc@1 67.97 (60.80) Acc@5 89.84 (84.38) - Test: [50/79] Time 0.077 (0.082) Acc@1 60.16 (60.71) Acc@5 87.50 (84.30) - Test: [60/79] Time 0.079 (0.082) Acc@1 67.19 (60.43) Acc@5 87.50 (83.72) - Test: [70/79] Time 0.082 (0.081) Acc@1 53.12 (60.17) Acc@5 80.47 (83.31) - * Acc@1 60.730 Acc@5 83.930 Total time: 6.353 - Accuracy of INT8 IR model: 60.730 + Test: [ 0/79] Time 0.189 (0.189) Acc@1 81.25 (81.25) Acc@5 91.41 (91.41) + Test: [10/79] Time 0.079 (0.091) Acc@1 59.38 (66.90) Acc@5 85.94 (87.43) + Test: [20/79] Time 0.078 (0.087) Acc@1 67.19 (64.25) Acc@5 85.16 (87.28) + Test: [30/79] Time 0.080 (0.085) Acc@1 51.56 (62.40) Acc@5 75.78 (85.21) + Test: [40/79] Time 0.077 (0.083) Acc@1 67.97 (60.94) Acc@5 89.84 (84.51) + Test: [50/79] Time 0.078 (0.082) Acc@1 62.50 (61.06) Acc@5 87.50 (84.45) + Test: [60/79] Time 0.081 (0.082) Acc@1 66.41 (60.71) Acc@5 85.94 (83.84) + Test: [70/79] Time 0.078 (0.082) Acc@1 52.34 (60.40) Acc@5 79.69 (83.42) + * Acc@1 60.930 Acc@5 84.020 Total time: 6.371 + Accuracy of INT8 IR model: 60.930 -IV. Compare performance of INT8 model and FP32 model in OpenVINO -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +IV. Compare performance of INT8 model and FP32 model in OpenVINO `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ Finally, measure the inference performance of the ``FP32`` and ``INT8`` models, using `Benchmark @@ -693,13 +721,13 @@ throughput (frames per second) values. .. parsed-literal:: Benchmark FP32 model (OpenVINO IR) - [ INFO ] Throughput: 37.61 FPS + [ INFO ] Throughput: 37.93 FPS Benchmark INT8 model (OpenVINO IR) - [ INFO ] Throughput: 159.33 FPS + [ INFO ] Throughput: 155.44 FPS Benchmark FP32 model (OpenVINO IR) synchronously - [ INFO ] Throughput: 39.06 FPS + [ INFO ] Throughput: 38.81 FPS Benchmark INT8 model (OpenVINO IR) synchronously - [ INFO ] Throughput: 140.24 FPS + [ INFO ] Throughput: 139.97 FPS Show device Information for reference: diff --git a/docs/notebooks/113-image-classification-quantization-with-output.rst b/docs/notebooks/113-image-classification-quantization-with-output.rst index 2cd7e3c7ed0..55f40d20ab8 100644 --- a/docs/notebooks/113-image-classification-quantization-with-output.rst +++ b/docs/notebooks/113-image-classification-quantization-with-output.rst @@ -1,10 +1,12 @@ Quantization of Image Classification Models =========================================== +.. _top: + This tutorial demonstrates how to apply ``INT8`` quantization to Image Classification model using `NNCF `__. It uses the -Mobilenet V2 model, trained on Cifar10 dataset. The code is designed to +MobileNet V2 model, trained on Cifar10 dataset. The code is designed to be extendable to custom models and datasets. The tutorial uses OpenVINO backend for performing model quantization in NNCF, if you interested how to apply quantization on PyTorch model, please check this @@ -12,12 +14,29 @@ to apply quantization on PyTorch model, please check this This tutorial consists of the following steps: -- Prepare the model for quantization. -- Define a data loading functionality. -- Perform quantization. -- Compare accuracy of the original and quantized models. -- Compare performance of the original and quantized models. -- Compare results on one picture. +- Prepare the model for quantization. +- Define a data loading functionality. +- Perform quantization. +- Compare accuracy of the original and quantized models. +- Compare performance of the original and quantized models. +- Compare results on one picture. + +**Table of contents**: + +- `Prepare the Model <#prepare-the-model>`__ +- `Prepare Dataset <#prepare-dataset>`__ +- `Perform Quantization <#perform-quantization>`__ + + - `Create Dataset for Validation <#create-dataset-for-validation>`__ + +- `Run nncf.quantize for Getting an Optimized Model <#run-nncf.quantize-for-getting-an-optimized-model>`__ +- `Serialize an OpenVINO IR model <#serialize-an-openvino-ir-model>`__ +- `Compare Accuracy of the Original and Quantized Models <#compare-accuracy-of-the-original-and-quantized-models>`__ + + - `Select inference device <#select-inference-device>`__ + +- `Compare Performance of the Original and Quantized Models <#compare-performance-of-the-original-and-quantized-models>`__ +- `Compare results on four pictures <#compare-results-on-four-pictures>`__ .. code:: ipython3 @@ -31,14 +50,16 @@ This tutorial consists of the following steps: DATA_DIR.mkdir(exist_ok=True) MODEL_DIR.mkdir(exist_ok=True) -Prepare the Model ------------------ +Prepare the Model `⇑ <#top>`__ +############################################################################################################################### + Model preparation stage has the following steps: -- Download a PyTorch model -- Convert model to OpenVINO Intermediate Representation format (IR) using Model Optimizer Python API -- Serialize converted model on disk +- Download a PyTorch model +- Convert model to OpenVINO Intermediate Representation format (IR) + using model conversion Python API +- Serialize converted model on disk .. code:: ipython3 @@ -57,7 +78,7 @@ Model preparation stage has the following steps: remote: Counting objects: 100% (281/281), done. remote: Compressing objects: 100% (95/95), done. remote: Total 282 (delta 136), reused 269 (delta 129), pack-reused 1 - Receiving objects: 100% (282/282), 9.22 MiB | 4.20 MiB/s, done. + Receiving objects: 100% (282/282), 9.22 MiB | 4.67 MiB/s, done. Resolving deltas: 100% (136/136), done. @@ -67,17 +88,17 @@ Model preparation stage has the following steps: model = cifar10_mobilenetv2_x1_0(pretrained=True) -OpenVINO support PyTorch models via conversion to OpenVINO Intermediate -Representation format. Model Optimizer Python API should be used for -conversion. ``mo.convert_model`` accept PyTorch model instance and -convert it into ``openvino.runtime.Model`` representation of model in -OpenVINO. Optionally, you may specify ``example_input`` which serves as -helper for model tracing and ``input_shape`` for conversion model with -static shape. Conveted model is ready to be loaded on device for -inference and can be saved on disk for next usage via ``serialize`` -function. More details about Model Optimizer Python API can be found on -this -`page `__. +OpenVINO supports PyTorch models via conversion to OpenVINO Intermediate +Representation format using model conversion Python API. +``mo.convert_model`` accept PyTorch model instance and convert it into +``openvino.runtime.Model`` representation of model in OpenVINO. +Optionally, you may specify ``example_input`` which serves as a helper +for model tracing and ``input_shape`` for converting the model with +static shape. The converted model is ready to be loaded on a device for +inference and can be saved on a disk for next usage via the +``serialize`` function. More details about model conversion Python API +can be found on this +`page `__. .. code:: ipython3 @@ -90,8 +111,9 @@ this serialize(ov_model, MODEL_DIR / "mobilenet_v2.xml") -Prepare Dataset ---------------- +Prepare Dataset `⇑ <#top>`__ +############################################################################################################################### + We will use `CIFAR10 `__ dataset from @@ -132,22 +154,24 @@ Preprocessing for model obtained from training Extracting ../data/datasets/cifar10/cifar-10-python.tar.gz to ../data/datasets/cifar10 -Perform Quantization --------------------- +Perform Quantization `⇑ <#top>`__ +############################################################################################################################### + `NNCF `__ provides a suite of advanced algorithms for Neural Networks inference optimization in OpenVINO with minimal accuracy drop. We will use 8-bit quantization in post-training mode (without the fine-tuning pipeline) to optimize -MobilenetV2. The optimization process contains the following steps: +MobileNetV2. The optimization process contains the following steps: 1. Create a Dataset for quantization. 2. Run ``nncf.quantize`` for getting an optimized model. 3. Serialize an OpenVINO IR model, using the ``openvino.runtime.serialize`` function. -Create Datset for Validation -~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Create Dataset for Validation `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + NNCF is compatible with ``torch.utils.data.DataLoader`` interface. For performing quantization it should be passed into ``nncf.Dataset`` object @@ -171,8 +195,9 @@ model during quantization, in our case, to pick input tensor from pair INFO:nncf:NNCF initialized successfully. Supported frameworks detected: torch, tensorflow, onnx, openvino -Run nncf.quantize for Getting an Optimized Model ------------------------------------------------- +Run nncf.quantize for Getting an Optimized Model `⇑ <#top>`__ +############################################################################################################################### + ``nncf.quantize`` function accepts model and prepared quantization dataset for performing basic quantization. Optionally, additional @@ -188,12 +213,13 @@ about supported parameters can be found on this .. parsed-literal:: - Statistics collection: 100%|██████████| 300/300 [00:08<00:00, 35.03it/s] - Biases correction: 100%|██████████| 36/36 [00:01<00:00, 23.98it/s] + Statistics collection: 100%|██████████| 300/300 [00:08<00:00, 35.62it/s] + Biases correction: 100%|██████████| 36/36 [00:01<00:00, 19.40it/s] -Serialize an OpenVINO IR model ------------------------------- +Serialize an OpenVINO IR model `⇑ <#top>`__ +############################################################################################################################### + Similar to ``mo.convert_model``, quantized model is ``openvino.runtime.Model`` object which ready to be loaded into device @@ -203,8 +229,9 @@ and can be serialized on disk using ``openvino.runtime.serialize``. serialize(quant_ov_model, MODEL_DIR / "quantized_mobilenet_v2.xml") -Compare Accuracy of the Original and Quantized Models ------------------------------------------------------ +Compare Accuracy of the Original and Quantized Models `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -221,10 +248,11 @@ Compare Accuracy of the Original and Quantized Models total += 1 return correct / total -Select inference device -~~~~~~~~~~~~~~~~~~~~~~~ +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -select device from dropdown list for running inference using OpenVINO + +Select device from dropdown list for running inference using OpenVINO: .. code:: ipython3 @@ -287,8 +315,9 @@ select device from dropdown list for running inference using OpenVINO Accuracy of the optimized model: 93.51% -Compare Performance of the Original and Quantized Models --------------------------------------------------------- +Compare Performance of the Original and Quantized Models `⇑ <#top>`__ +############################################################################################################################### + Finally, measure the inference performance of the ``FP32`` and ``INT8`` models, using `Benchmark @@ -328,18 +357,18 @@ Tool `__ +############################################################################################################################### + .. code:: ipython3 @@ -551,5 +581,5 @@ Compare results on four pictures -.. image:: 113-image-classification-quantization-with-output_files/113-image-classification-quantization-with-output_28_2.png +.. image:: 113-image-classification-quantization-with-output_files/113-image-classification-quantization-with-output_29_2.png diff --git a/docs/notebooks/113-image-classification-quantization-with-output_files/113-image-classification-quantization-with-output_28_2.png b/docs/notebooks/113-image-classification-quantization-with-output_files/113-image-classification-quantization-with-output_29_2.png similarity index 100% rename from docs/notebooks/113-image-classification-quantization-with-output_files/113-image-classification-quantization-with-output_28_2.png rename to docs/notebooks/113-image-classification-quantization-with-output_files/113-image-classification-quantization-with-output_29_2.png diff --git a/docs/notebooks/113-image-classification-quantization-with-output_files/index.html b/docs/notebooks/113-image-classification-quantization-with-output_files/index.html index 99d7ae56219..c510054b810 100644 --- a/docs/notebooks/113-image-classification-quantization-with-output_files/index.html +++ b/docs/notebooks/113-image-classification-quantization-with-output_files/index.html @@ -1,7 +1,7 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/113-image-classification-quantization-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/113-image-classification-quantization-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/113-image-classification-quantization-with-output_files/


../
-113-image-classification-quantization-with-outp..> 12-Jul-2023 00:11               14855
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/113-image-classification-quantization-with-output_files/


../
+113-image-classification-quantization-with-outp..> 16-Aug-2023 01:31               14855
 

diff --git a/docs/notebooks/115-async-api-with-output.rst b/docs/notebooks/115-async-api-with-output.rst index c7bee23e5b9..4daffcb8ab8 100644 --- a/docs/notebooks/115-async-api-with-output.rst +++ b/docs/notebooks/115-async-api-with-output.rst @@ -1,8 +1,10 @@ Asynchronous Inference with OpenVINO™ ===================================== +.. _top: + This notebook demonstrates how to use the `Async -API `__ +API `__ for asynchronous execution with OpenVINO. OpenVINO Runtime supports inference in either synchronous or @@ -11,12 +13,37 @@ device is busy with inference, the application can perform other tasks in parallel (for example, populating inputs or scheduling other requests) rather than wait for the current inference to complete first. -Imports -------- + +**Table of contents**: + +- `Imports <#imports>`__ +- `Prepare model and data processing <#prepare-model-and-data-processing>`__ + + - `Download test model <#download-test-model>`__ + - `Load the model <#load-the-model>`__ + - `Create functions for data processing <#create-functions-for-data-processing>`__ + - `Get the test video <#get-the-test-video>`__ + +- `How to improve the throughput of video processing <#how-to-improve-the-throughput-of-video-processing>`__ + + - `Sync Mode (default) <#sync-mode-default>`__ + - `Test performance in Sync Mode <#test-performance-in-sync-mode>`__ + - `Async Mode <#async-mode>`__ + - `Test the performance in Async Mode <#test-the-performance-in-async-mode>`__ + - `Compare the performance <#compare-the-performance>`__ + +- `AsyncInferQueue <#asyncinferqueue>`__ + + - `Setting Callback <#setting-callback>`__ + - `Test the performance with AsyncInferQueue <#test-the-performance-with-asyncinferqueue>`__ + +Imports `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 - !pip install -q 'openvino-dev>=2023.0.0' + !pip install -q "openvino-dev>=2023.0.0" !pip install -q opencv-python matplotlib .. code:: ipython3 @@ -38,14 +65,15 @@ Imports import notebook_utils as utils -Prepare model and data processing ---------------------------------- +Prepare model and data processing `⇑ <#top>`__ +############################################################################################################################### -Download test model -~~~~~~~~~~~~~~~~~~~ -We use a pre-trained model from OpenVINO’s `Open Model -Zoo `__ to start the +Download test model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + +We use a pre-trained model from OpenVINO’s +`Open Model Zoo `__ to start the test. In this case, the model will be executed to detect the person in each frame of the video. @@ -80,8 +108,9 @@ each frame of the video. -Load the model -~~~~~~~~~~~~~~ +Load the model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -91,7 +120,7 @@ Load the model # read the network and corresponding weights from file model = ie.read_model(model=model_path) - # compile the model for the CPU (you can choose manually CPU, GPU, etc.) + # compile the model for the CPU (you can choose manually CPU, GPU etc.) # or let the engine choose the best available device (AUTO) compiled_model = ie.compile_model(model=model, device_name="CPU") @@ -100,8 +129,9 @@ Load the model N, C, H, W = input_layer_ir.shape shape = (H, W) -Create functions for data processing -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Create functions for data processing `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -142,27 +172,30 @@ Create functions for data processing cv2.putText(image, str(round(fps, 2)) + " fps", (5, 20), cv2.FONT_HERSHEY_SIMPLEX, 0.7, (0, 255, 0), 3) return image -Get the test video -~~~~~~~~~~~~~~~~~~ +Get the test video `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 video_path = 'https://storage.openvinotoolkit.org/repositories/openvino_notebooks/data/data/video/CEO%20Pat%20Gelsinger%20on%20Leading%20Intel.mp4' -How to improve the throughput of video processing -------------------------------------------------- +How to improve the throughput of video processing `⇑ <#top>`__ +############################################################################################################################### Below, we compare the performance of the synchronous and async-based approaches: -Sync Mode (default) -~~~~~~~~~~~~~~~~~~~ +Sync Mode (default) `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Let us see how video processing works with the default approach. Using -the synchronous approach, the frame is captured with OpenCV and then -immediately processed: +Let us see how video processing works with the default approach. Using the synchronous approach, the frame is +captured with OpenCV and then immediately processed: -.. image:: https://camo.githubusercontent.com/b77ae49e3c46a0fe3931ad02a7767d6def49aa7b2a7a8ad620a4e12ae31d43a5/68747470733a2f2f757365722d696d616765732e67697468756275736572636f6e74656e742e636f6d2f39313233373932342f3136383435323537332d64333534656135622d373936362d343465352d383133642d6639303533626534333338612e706e67 +.. figure:: https://user-images.githubusercontent.com/91237924/168452573-d354ea5b-7966-44e5-813d-f9053be4338a.png + :alt: drawing + + drawing :: @@ -174,6 +207,8 @@ immediately processed: // display CURRENT result } +\``\` + .. code:: ipython3 def sync_api(source, flip, fps, use_popup, skip_first_frames): @@ -240,8 +275,9 @@ immediately processed: player.stop() return sync_fps -Test performance in Sync Mode -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Test performance in Sync Mode `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -250,15 +286,18 @@ Test performance in Sync Mode +.. image:: 115-async-api-with-output_files/115-async-api-with-output_15_0.png + .. parsed-literal:: Source ended - average throuput in sync mode: 38.25 fps + average throuput in sync mode: 37.71 fps -Async Mode -~~~~~~~~~~ +Async Mode `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Let us see how the OpenVINO Async API can improve the overall frame rate of an application. The key advantage of the Async approach is as @@ -267,7 +306,10 @@ do other things in parallel (for example, populating inputs or scheduling other requests) rather than wait for the current inference to complete first. -.. image:: https://camo.githubusercontent.com/7bcadb7cf72aefc84d74d8e971a31dab1710bb4303fef814619df638dff8be92/68747470733a2f2f757365722d696d616765732e67697468756275736572636f6e74656e742e636f6d2f39313233373932342f3136383435323537322d63326666316335392d643437302d346238352d623166362d6236653164616339353430652e706e67 +.. figure:: https://user-images.githubusercontent.com/91237924/168452572-c2ff1c59-d470-4b85-b1f6-b6e1dac9540e.png + :alt: drawing + + drawing In the example below, inference is applied to the results of the video decoding. So it is possible to keep multiple infer requests, and while @@ -369,8 +411,9 @@ pipeline (decoding vs inference) and not by the sum of the stages. player.stop() return async_fps -Test the performance in Async Mode -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Test the performance in Async Mode `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -378,14 +421,19 @@ Test the performance in Async Mode print(f"average throuput in async mode: {async_fps:.2f} fps") + +.. image:: 115-async-api-with-output_files/115-async-api-with-output_19_0.png + + .. parsed-literal:: Source ended - average throuput in async mode: 71.88 fps + average throuput in async mode: 73.36 fps -Compare the performance -~~~~~~~~~~~~~~~~~~~~~~~ +Compare the performance `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -409,24 +457,29 @@ Compare the performance -AsyncInferQueue ---------------- +.. image:: 115-async-api-with-output_files/115-async-api-with-output_21_0.png + + +``AsyncInferQueue`` `⇑ <#top>`__ +############################################################################################################################### + Asynchronous mode pipelines can be supported with the `AsyncInferQueue `__ -wrapper class. This class automatically spawns the pool of InferRequest -objects (also called “jobs”) and provides synchronization mechanisms to -control the flow of the pipeline. It is a simpler way to manage the -infer request queue in Asynchronous mode. +wrapper class. This class automatically spawns the pool of +``InferRequest`` objects (also called “jobs”) and provides +synchronization mechanisms to control the flow of the pipeline. It is a +simpler way to manage the infer request queue in Asynchronous mode. + +Setting Callback `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Setting Callback -~~~~~~~~~~~~~~~~ When ``callback`` is set, any job that ends inference calls upon the Python function. The ``callback`` function must have two arguments: one is the request that calls the ``callback``, which provides the -InferRequest API; the other is called “userdata”, which provides the -possibility of passing runtime values. +``InferRequest`` API; the other is called “user data”, which provides +the possibility of passing runtime values. .. code:: ipython3 @@ -496,8 +549,9 @@ possibility of passing runtime values. infer_queue.wait_all() player.stop() -Test the performance with AsyncInferQueue -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Test the performance with ``AsyncInferQueue`` `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -508,7 +562,10 @@ Test the performance with AsyncInferQueue +.. image:: 115-async-api-with-output_files/115-async-api-with-output_27_0.png + + .. parsed-literal:: - average throughput in async mode with async infer queue: 104.69 fps + average throughput in async mode with async infer queue: 103.73 fps diff --git a/docs/notebooks/115-async-api-with-output_files/115-async-api-with-output_21_0.png b/docs/notebooks/115-async-api-with-output_files/115-async-api-with-output_21_0.png new file mode 100644 index 00000000000..60870d8f525 --- /dev/null +++ b/docs/notebooks/115-async-api-with-output_files/115-async-api-with-output_21_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:b94a4331de267d1151abf422719316d87b59282125c1cb8471a99e34d561bdd1 +size 30455 diff --git a/docs/notebooks/115-async-api-with-output_files/index.html b/docs/notebooks/115-async-api-with-output_files/index.html index 70deb4ff8a8..9dad300ecd4 100644 --- a/docs/notebooks/115-async-api-with-output_files/index.html +++ b/docs/notebooks/115-async-api-with-output_files/index.html @@ -1,10 +1,10 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/115-async-api-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/115-async-api-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/115-async-api-with-output_files/


../
-115-async-api-with-output_15_0.png                 12-Jul-2023 00:11                4307
-115-async-api-with-output_19_0.png                 12-Jul-2023 00:11                4307
-115-async-api-with-output_21_0.png                 12-Jul-2023 00:11               30421
-115-async-api-with-output_27_0.png                 12-Jul-2023 00:11                4307
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/115-async-api-with-output_files/


../
+115-async-api-with-output_15_0.png                 16-Aug-2023 01:31                4307
+115-async-api-with-output_19_0.png                 16-Aug-2023 01:31                4307
+115-async-api-with-output_21_0.png                 16-Aug-2023 01:31               30455
+115-async-api-with-output_27_0.png                 16-Aug-2023 01:31                4307
 

diff --git a/docs/notebooks/116-sparsity-optimization-with-output.rst b/docs/notebooks/116-sparsity-optimization-with-output.rst index f61fe0c191a..aa321a6b57e 100644 --- a/docs/notebooks/116-sparsity-optimization-with-output.rst +++ b/docs/notebooks/116-sparsity-optimization-with-output.rst @@ -1,14 +1,14 @@ Accelerate Inference of Sparse Transformer Models with OpenVINO™ and 4th Gen Intel® Xeon® Scalable Processors ============================================================================================================= +.. _top: + This tutorial demonstrates how to improve performance of sparse Transformer models with `OpenVINO `__ on 4th Gen Intel® Xeon® Scalable processors. -The tutorial downloads `a BERT-base -model `__ -which has been quantized, sparsified, and tuned for `SST2 -datasets `__ using +The tutorial downloads `a BERT-base model `__ +which has been quantized, sparsified, and tuned for `SST2 datasets `__ using `Optimum-Intel `__. It demonstrates the inference performance advantage on 4th Gen Intel® Xeon® Scalable Processors by running it with `Sparse Weight @@ -21,16 +21,29 @@ consists of the following steps: integration with Hugging Face Optimum. - Compare sparse 8-bit vs. dense 8-bit inference performance. -Prerequisites -------------- +**Table of contents**: + +- `Prerequisites <#prerequisites>`__ +- `Imports <#imports>`__ + + - `Download, quantize and sparsify the model, using Hugging Face Optimum API <#download-quantize-and-sparsify-the-model-using-hugging-face-optimum-api>`__ + +- `Benchmark quantized dense inference performance <#benchmark-quantized-dense-inference-performance>`__ +- `Benchmark quantized sparse inference performance <#benchmark-quantized-sparse-inference-performance>`__ +- `When this might be helpful <#when-this-might-be-helpful>`__ + +Prerequisites `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 - !pip install -q 'openvino-dev>=2023.0.0' + !pip install -q "openvino-dev>=2023.0.0" !pip install -q "git+https://github.com/huggingface/optimum-intel.git" datasets onnx onnxruntime -Imports -------- +Imports `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -44,10 +57,10 @@ Imports .. parsed-literal:: - 2023-07-11 22:51:18.663091: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. - 2023-07-11 22:51:18.697477: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + 2023-08-15 22:55:04.775263: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-08-15 22:55:04.809127: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. - 2023-07-11 22:51:19.248196: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + 2023-08-15 22:55:05.351203: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT .. parsed-literal:: @@ -60,8 +73,8 @@ Imports No CUDA runtime is found, using CUDA_HOME='/usr/local/cuda' -Download, quantize and sparsify the model, using Hugging Face Optimum API -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Download, quantize and sparsify the model, using Hugging Face Optimum API. `⇑ <#top>`__ +############################################################################################################################### The first step is to download a quantized sparse transformers which has been translated to OpenVINO IR. Then, it will be put through a @@ -128,14 +141,14 @@ the IRs into a single folder. -Benchmark quantized dense inference performance ------------------------------------------------ +Benchmark quantized dense inference performance `⇑ <#top>`__ +############################################################################################################################### -Benchmark dense inference performance using parallel execution on four -CPU cores to simulate a small instance in the cloud infrastructure. -Sequence length is dependent on use cases, 16 is common for -conversational AI while 160 for question answering task. It is set to 64 -as an example. It is recommended to tune based on your applications. +Benchmark dense inference performance using parallel execution on four CPU cores +to simulate a small instance in the cloud infrastructure. Sequence +length is dependent on use cases, 16 is common for conversational AI +while 160 for question answering task. It is set to 64 as an example. It +is recommended to tune based on your applications. .. code:: ipython3 @@ -175,7 +188,7 @@ as an example. It is recommended to tune based on your applications. [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.THROUGHPUT. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 79.41 ms + [ INFO ] Read model took 74.55 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] input_ids (node: input_ids) : i64 / [...] / [?,?] @@ -186,7 +199,7 @@ as an example. It is recommended to tune based on your applications. [Step 5/11] Resizing model to match image sizes and given batch [ INFO ] Model batch size: 1 [ INFO ] Reshaping model: 'input_ids': [1,64], 'attention_mask': [1,64], 'token_type_ids': [1,64] - [ INFO ] Reshape model took 26.38 ms + [ INFO ] Reshape model took 26.03 ms [Step 6/11] Configuring input of the model [ INFO ] Model inputs: [ INFO ] input_ids (node: input_ids) : i64 / [...] / [1,64] @@ -195,7 +208,7 @@ as an example. It is recommended to tune based on your applications. [ INFO ] Model outputs: [ INFO ] logits (node: logits) : f32 / [...] / [1,2] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 1252.95 ms + [ INFO ] Compile model took 1231.43 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: [ INFO ] NETWORK_NAME: torch_jit @@ -221,21 +234,22 @@ as an example. It is recommended to tune based on your applications. [ INFO ] Fill input 'token_type_ids' with random values [Step 10/11] Measuring performance (Start inference asynchronously, 4 inference requests, limits: 60000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 26.89 ms + [ INFO ] First inference took 31.02 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 8944 iterations - [ INFO ] Duration: 60031.11 ms + [ INFO ] Count: 8896 iterations + [ INFO ] Duration: 60044.19 ms [ INFO ] Latency: - [ INFO ] Median: 26.64 ms - [ INFO ] Average: 26.68 ms - [ INFO ] Min: 25.23 ms - [ INFO ] Max: 40.15 ms - [ INFO ] Throughput: 148.99 FPS + [ INFO ] Median: 26.82 ms + [ INFO ] Average: 26.87 ms + [ INFO ] Min: 25.13 ms + [ INFO ] Max: 38.04 ms + [ INFO ] Throughput: 148.16 FPS -Benchmark quantized sparse inference performance ------------------------------------------------- +Benchmark quantized sparse inference performance `⇑ <#top>`__ +############################################################################################################################### + To enable sparse weight decompression feature, users can add it to runtime config like below. ``CPU_SPARSE_WEIGHTS_DECOMPRESSION_RATE`` @@ -282,7 +296,7 @@ for which a layer will be enabled. [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.THROUGHPUT. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 69.37 ms + [ INFO ] Read model took 63.96 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] input_ids (node: input_ids) : i64 / [...] / [?,?] @@ -293,7 +307,7 @@ for which a layer will be enabled. [Step 5/11] Resizing model to match image sizes and given batch [ INFO ] Model batch size: 1 [ INFO ] Reshaping model: 'input_ids': [1,64], 'attention_mask': [1,64], 'token_type_ids': [1,64] - [ INFO ] Reshape model took 25.96 ms + [ INFO ] Reshape model took 26.17 ms [Step 6/11] Configuring input of the model [ INFO ] Model inputs: [ INFO ] input_ids (node: input_ids) : i64 / [...] / [1,64] @@ -302,7 +316,7 @@ for which a layer will be enabled. [ INFO ] Model outputs: [ INFO ] logits (node: logits) : f32 / [...] / [1,2] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 1250.91 ms + [ INFO ] Compile model took 1252.94 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: [ INFO ] NETWORK_NAME: torch_jit @@ -328,21 +342,22 @@ for which a layer will be enabled. [ INFO ] Fill input 'token_type_ids' with random values [Step 10/11] Measuring performance (Start inference asynchronously, 4 inference requests, limits: 60000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 32.79 ms + [ INFO ] First inference took 30.12 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 8972 iterations - [ INFO ] Duration: 60029.38 ms + [ INFO ] Count: 8840 iterations + [ INFO ] Duration: 60036.27 ms [ INFO ] Latency: - [ INFO ] Median: 26.59 ms - [ INFO ] Average: 26.63 ms - [ INFO ] Min: 25.42 ms - [ INFO ] Max: 41.12 ms - [ INFO ] Throughput: 149.46 FPS + [ INFO ] Median: 26.83 ms + [ INFO ] Average: 26.88 ms + [ INFO ] Min: 26.11 ms + [ INFO ] Max: 40.45 ms + [ INFO ] Throughput: 147.24 FPS -When this might be helpful --------------------------- +When this might be helpful `⇑ <#top>`__ +############################################################################################################################### + This feature can improve inference performance for models with sparse weights in the scenarios when the model is deployed to handle multiple diff --git a/docs/notebooks/117-model-server-with-output.rst b/docs/notebooks/117-model-server-with-output.rst index 91658f0d20c..54989d2a0e7 100644 --- a/docs/notebooks/117-model-server-with-output.rst +++ b/docs/notebooks/117-model-server-with-output.rst @@ -1,6 +1,8 @@ Hello Model Server ================== +.. _top: + Introduction to OpenVINO™ Model Server (OVMS). What is Model Serving? @@ -12,37 +14,54 @@ the model server, which performs inference and sends a response back to the client. Model serving offers many advantages for efficient model deployment: -- Remote inference enables using lightweight clients with only the - necessary functions to perform API calls to edge or cloud - deployments. -- Applications are independent of the model framework, hardware device, - and infrastructure. -- Client applications in any programming language that supports REST or - gRPC calls can be used to run inference remotely on the model server. -- Clients require fewer updates since client libraries change very - rarely. -- Model topology and weights are not exposed directly to client - applications, making it easier to control access to the model. -- Ideal architecture for microservices-based applications and - deployments in cloud environments – including Kubernetes and - OpenShift clusters. -- Efficient resource utilization with horizontal and vertical inference - scaling. +- Remote inference enables using lightweight clients with only the + necessary functions to perform API calls to edge or cloud + deployments. +- Applications are independent of the model framework, hardware device, + and infrastructure. +- Client applications in any programming language that supports REST or + gRPC calls can be used to run inference remotely on the model server. +- Clients require fewer updates since client libraries change very + rarely. +- Model topology and weights are not exposed directly to client + applications, making it easier to control access to the model. +- Ideal architecture for microservices-based applications and + deployments in cloud environments – including Kubernetes and + OpenShift clusters. +- Efficient resource utilization with horizontal and vertical inference + scaling. -.. figure:: https://user-images.githubusercontent.com/91237924/215658773-4720df00-3b95-4a84-85a2-40f06138e914.png - :alt: ovms_diagram +|ovms_diagram| - ovms_diagram +**Table of contents**: -Serving with OpenVINO Model Server ----------------------------------- +- `Serving with OpenVINO Model Server <#serving-with-openvino-model-server1>`__ +- `Step 1: Prepare Docker <#step-1-prepare-docker>`__ +- `Step 2: Preparing a Model Repository <#step-2-preparing-a-model-repository>`__ +- `Step 3: Start the Model Server Container <#start-the-model-server-container>`__ +- `Step 4: Prepare the Example Client Components <#prepare-the-example-client-components>`__ -OpenVINO Model Server (OVMS) is a high-performance system for serving -models. Implemented in C++ for scalability and optimized for deployment -on Intel architectures, the model server uses the same architecture and -API as TensorFlow Serving and KServe while applying OpenVINO for -inference execution. Inference service is provided via gRPC or REST API, -making deploying new algorithms and AI experiments easy. + - `Prerequisites <#prerequisites>`__ + - `Imports <#imports>`__ + - `Request Model Status <#request-model-status>`__ + - `Request Model Metadata <#request-model-metadata>`__ + - `Load input image <#load-input-image>`__ + - `Request Prediction on a Numpy Array <#request-prediction-on-a-numpy-array>`__ + - `Visualization <#visualization>`__ + +- `References <#references>`__ + +.. |ovms_diagram| image:: https://user-images.githubusercontent.com/91237924/215658773-4720df00-3b95-4a84-85a2-40f06138e914.png + +Serving with OpenVINO Model Server `⇑ <#top>`__ +############################################################################################################################### + +OpenVINO Model Server (OVMS) is a high-performance system for serving models. Implemented in +C++ for scalability and optimized for deployment on Intel architectures, +the model server uses the same architecture and API as TensorFlow +Serving and KServe while applying OpenVINO for inference execution. +Inference service is provided via gRPC or REST API, making deploying new +algorithms and AI experiments easy. .. figure:: https://user-images.githubusercontent.com/91237924/215658767-0e0fc221-aed0-4db1-9a82-6be55f244dba.png :alt: ovms_high_level @@ -51,11 +70,10 @@ making deploying new algorithms and AI experiments easy. To quickly start using OpenVINO™ Model Server, follow these steps: -Step 1: Prepare Docker ----------------------- +Step 1: Prepare Docker `⇑ <#top>`__ +############################################################################################################################### -Install `Docker Engine `__, -including its +Install `Docker Engine `__, including its `post-installation `__ steps, on your development system. To verify installation, test it, using the following command. When it is ready, it will display a test @@ -92,11 +110,11 @@ image and a message. -Step 2: Preparing a Model Repository ------------------------------------- +Step 2: Preparing a Model Repository `⇑ <#top>`__ +############################################################################################################################### -The models need to be placed and mounted in a particular directory -structure and according to the following rules: +The models need to be placed and mounted in a particular directory structure and according to +the following rules: :: @@ -162,11 +180,13 @@ structure and according to the following rules: model_bin_url = "https://storage.openvinotoolkit.org/repositories/open_model_zoo/2022.3/models_bin/1/horizontal-text-detection-0001/FP32/horizontal-text-detection-0001.bin" download_file(model_xml_url, XML_PATH, MODEL_DIR) - download_file(model_bin_url, BIN_PATH_name, MODEL_DIR)model_xml_url = "https://storage.openvinotoolkit.org/repositories/open_model_zoo/2022.3/models_bin/1/horizontal-text-detection-0001/FP32/horizontal-text-detection-0001.xml" - model_bin_url = "https://storage.openvinotoolkit.org/repositories/open_model_zoo/2022.3/models_bin/1/horizontal-text-detection-0001/FP32/horizontal-text-detection-0001.bin" + download_file(model_bin_url, BIN_PATH_name, MODEL_DIR) - download_file(model_xml_url, model_xml_name, base_model_dir) - download_file(model_bin_url, model_bin_name, base_model_dir) + model_xml_url = "https://storage.openvinotoolkit.org/repositories/open_model_zoo/2022.3/models_bin/1/horizontal-text-detection-0001/FP32/horizontal-text-detection-0001.xml" + model_bin_url = "https://storage.openvinotoolkit.org/repositories/open_model_zoo/2022.3/models_bin/1/horizontal-text-detection-0001/FP32/horizontal-text-detection-0001.bin" + + download_file(model_xml_url, model_xml_name, base_model_dir) + download_file(model_bin_url, model_bin_name, base_model_dir) .. parsed-literal:: @@ -174,8 +194,8 @@ structure and according to the following rules: Model Copied to "./models/detection/1". -Step 3: Start the Model Server Container ----------------------------------------- +Step 3: Start the Model Server Container `⇑ <#top>`__ +############################################################################################################################### Pull and start the container: @@ -202,8 +222,8 @@ Check whether the OVMS container is running normally: The required Model Server parameters are listed below. For additional -configuration options, see the `Model Server Parameters -section `__. +configuration options, see the +`Model Server Parameters section `__. .. raw:: html @@ -419,7 +439,7 @@ openvino/model_server:latest .. container:: line - represents the image name; the ovms binary is the Docker entry + represents the image name; the OVMS binary is the Docker entry point .. container:: line @@ -623,8 +643,8 @@ openvino/model_server:latest If the serving port ``9000`` is already in use, please switch it to another available port on your system. For example:\ ``-p 9020:9000`` -Step 4: Prepare the Example Client Components ---------------------------------------------- +Step 4: Prepare the Example Client Components `⇑ <#top>`__ +############################################################################################################################### OpenVINO Model Server exposes two sets of APIs: one compatible with ``TensorFlow Serving`` and another one, with ``KServe API``, for @@ -634,8 +654,9 @@ into existing systems the already leverage one of these APIs for inference. This example will demonstrate how to write a TensorFlow Serving API client for object detection. -Prerequisites -~~~~~~~~~~~~~ +Prerequisites `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Install necessary packages. @@ -669,8 +690,9 @@ Install necessary packages. You should consider upgrading via the '/home/adrian/repos/openvino_notebooks_adrian/venv/bin/python -m pip install --upgrade pip' command. -Imports -~~~~~~~ +Imports `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -679,8 +701,9 @@ Imports import matplotlib.pyplot as plt from ovmsclient import make_grpc_client -Request Model Status -~~~~~~~~~~~~~~~~~~~~ +Request Model Status `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -697,8 +720,9 @@ Request Model Status {1: {'state': 'AVAILABLE', 'error_code': 0, 'error_message': 'OK'}} -Request Model Metadata -~~~~~~~~~~~~~~~~~~~~~~ +Request Model Metadata `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -711,8 +735,9 @@ Request Model Metadata {'model_version': 1, 'inputs': {'image': {'shape': [1, 3, 704, 704], 'dtype': 'DT_FLOAT'}}, 'outputs': {'1469_1470.0': {'shape': [-1], 'dtype': 'DT_FLOAT'}, '1078_1079.0': {'shape': [1000], 'dtype': 'DT_FLOAT'}, '1330_1331.0': {'shape': [36], 'dtype': 'DT_FLOAT'}, 'labels': {'shape': [-1], 'dtype': 'DT_INT32'}, '1267_1268.0': {'shape': [121], 'dtype': 'DT_FLOAT'}, '1141_1142.0': {'shape': [1000], 'dtype': 'DT_FLOAT'}, '1204_1205.0': {'shape': [484], 'dtype': 'DT_FLOAT'}, 'boxes': {'shape': [-1, 5], 'dtype': 'DT_FLOAT'}}} -Load input image -~~~~~~~~~~~~~~~~ +Load input image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -742,8 +767,9 @@ Load input image .. image:: 117-model-server-with-output_files/117-model-server-with-output_20_1.png -Request Prediction on a Numpy Array -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Request Prediction on a Numpy Array `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -767,8 +793,9 @@ Request Prediction on a Numpy Array [2.2261986e+01 4.5406548e+01 1.8868817e+02 1.0225631e+02 3.0407205e-01]] -Visualization -~~~~~~~~~~~~~ +Visualization `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -850,9 +877,11 @@ command: ovms -References ----------- +References `⇑ <#top>`__ +############################################################################################################################### -1. `OpenVINO™ Model - Server `__ -2. `openvinotoolkit/model_server `__ + +1. `OpenVINO™ Model Server + documentation `__ +2. `OpenVINO™ Model Server GitHub + repository `__ diff --git a/docs/notebooks/117-model-server-with-output_files/index.html b/docs/notebooks/117-model-server-with-output_files/index.html index 12d8159dd5d..bc4151c4fbe 100644 --- a/docs/notebooks/117-model-server-with-output_files/index.html +++ b/docs/notebooks/117-model-server-with-output_files/index.html @@ -1,8 +1,8 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/117-model-server-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/117-model-server-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/117-model-server-with-output_files/


../
-117-model-server-with-output_20_1.png              12-Jul-2023 00:11              112408
-117-model-server-with-output_25_1.png              12-Jul-2023 00:11              232667
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/117-model-server-with-output_files/


../
+117-model-server-with-output_20_1.png              16-Aug-2023 01:31              112408
+117-model-server-with-output_25_1.png              16-Aug-2023 01:31              232667
 

diff --git a/docs/notebooks/118-optimize-preprocessing-with-output.rst b/docs/notebooks/118-optimize-preprocessing-with-output.rst index d3d56233839..c76a8986137 100644 --- a/docs/notebooks/118-optimize-preprocessing-with-output.rst +++ b/docs/notebooks/118-optimize-preprocessing-with-output.rst @@ -1,6 +1,8 @@ Optimize Preprocessing ====================== +.. _top: + When input data does not fit the model input tensor perfectly, additional operations/steps are needed to transform the data to the format expected by the model. This tutorial demonstrates how it could be @@ -15,18 +17,59 @@ and This tutorial include following steps: -- Downloading the model. -- Setup preprocessing with ModelOptimizer, loading the model and inference with original image. -- Setup preprocessing with Preprocessing API, loading the model and inference with original image. -- Fitting image to the model input type and inference with prepared image. -- Comparing results on one picture. -- Comparing performance. +- Downloading the model. +- Setup preprocessing with model conversion API, loading the model and + inference with original image. +- Setup preprocessing with Preprocessing API, loading the model and + inference with original image. +- Fitting image to the model input type and inference with prepared + image. +- Comparing results on one picture. +- Comparing performance. -Settings --------- +**Table of contents**: + +- `Settings <#settings>`__ +- `Imports <#imports>`__ + + - `Setup image and device <#setup-image-and-device>`__ + - `Downloading the model <#downloading-the-model>`__ + - `Create core <#create-core>`__ + - `Check the original parameters of image <#check-the-original-parameters-of-image>`__ + +- `Convert model to OpenVINO IR and setup preprocessing steps with model conversion API <#convert-model-to-openvino-ir-and-setup-preprocessing-steps-with-model-conversion-api>`__ + + - `Prepare image <#prepare-image>`__ + - `Compile model and perform inference <#compile-model-and-perform-inference>`__ + +- `Setup preprocessing steps with Preprocessing API and perform inference <#setup-preprocessing-steps-with-preprocessing-api-and-perform-inference>`__ + + - `Convert model to OpenVINO IR with model conversion API <#convert-model-to-openvino-ir-with-model-conversion-api>`__ + - `Create PrePostProcessor Object <#create-prepostprocessor-object>`__ + - `Declare User’s Data Format <#declare-users-data-format>`__ + - `Declaring Model Layout <#declaring-model-layout>`__ + - `Preprocessing Steps <#preprocessing-steps>`__ + - `Integrating Steps into a Model <#integrating-steps-into-a-model>`__ + +- `Load model and perform inference <#load-model-and-perform-inference>`__ +- `Fit image manually and perform inference <#fit-image-manually-and-perform-inference>`__ + + - `Load the model <#load-the-model>`__ + - `Load image and fit it to model input <#load-image-and-fit-it-to-model-input>`__ + - `Perform inference <#perform-inference>`__ + +- `Compare results <#compare-results>`__ + + - `Compare results on one image <#compare-results-on-one-image>`__ + - `Compare performance <#compare-performance>`__ + +Settings `⇑ <#top>`__ +############################################################################################################################### + + +Imports `⇑ <#top>`__ +############################################################################################################################### -Imports -------- .. code:: ipython3 @@ -43,14 +86,15 @@ Imports .. parsed-literal:: - 2023-07-11 22:53:33.947684: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. - 2023-07-11 22:53:33.981920: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + 2023-08-15 22:57:19.952994: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-08-15 22:57:19.987688: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. - 2023-07-11 22:53:34.528705: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + 2023-08-15 22:57:20.520711: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT -Setup image and device -~~~~~~~~~~~~~~~~~~~~~~ +Setup image and device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -79,8 +123,9 @@ Setup image and device -Downloading the model -~~~~~~~~~~~~~~~~~~~~~ +Downloading the model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + This tutorial uses the `InceptionResNetV2 `__. @@ -111,7 +156,7 @@ and save it to the disk. .. parsed-literal:: - 2023-07-11 22:53:35.902963: W tensorflow/core/common_runtime/gpu/gpu_device.cc:1956] Cannot dlopen some GPU libraries. Please make sure the missing libraries mentioned above are installed properly if you would like to use GPU. Follow the guide at https://www.tensorflow.org/install/gpu for how to download and setup the required libraries for your platform. + 2023-08-15 22:57:21.888060: W tensorflow/core/common_runtime/gpu/gpu_device.cc:1956] Cannot dlopen some GPU libraries. Please make sure the missing libraries mentioned above are installed properly if you would like to use GPU. Follow the guide at https://www.tensorflow.org/install/gpu for how to download and setup the required libraries for your platform. Skipping registering GPU devices... @@ -135,15 +180,17 @@ and save it to the disk. INFO:tensorflow:Assets written to: model/InceptionResNetV2/assets -Create core -~~~~~~~~~~~ +Create core `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 core = Core() -Check the original parameters of image -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Check the original parameters of image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -160,16 +207,15 @@ Check the original parameters of image -.. image:: 118-optimize-preprocessing-with-output_files/118-optimize-preprocessing-with-output_12_1.png +.. image:: 118-optimize-preprocessing-with-output_files/118-optimize-preprocessing-with-output_13_1.png -Convert model to OpenVINO IR and setup preprocessing steps with Model Optimizer -------------------------------------------------------------------------------- +Convert model to OpenVINO IR and setup preprocessing steps with model conversion API. `⇑ <#top>`__ +############################################################################################################################### -Use Model Optimizer to convert a TensorFlow model to OpenVINO IR. -``mo.convert_model`` python function will be used for converting model -using `OpenVINO Model -Optimizer `__. +To convert a TensorFlow model to OpenVINO IR, use the +``mo.convert_model`` python function of `model conversion +API `__. The function returns instance of OpenVINO Model class, which is ready to use in Python interface but can also be serialized to OpenVINO IR format for future execution using ``openvino.runtime.serialize``. The models @@ -182,9 +228,12 @@ pre-processing sub-graphs into the converted model. Setup the following conversions: -- mean normalization with ``mean_values`` parameter -- scale with ``scale_values`` -- color conversion, the color format of example image will be ``BGR``, but the model required ``RGB`` format, so add ``reverse_input_channels=True`` to process the image into the desired format +- mean normalization with ``mean_values`` parameter. +- scale with ``scale_values``. +- color conversion, the color format of example image will be ``BGR``, + but the model required ``RGB`` format, so add + ``reverse_input_channels=True`` to process the image into the desired + format. Also converting of layout could be specified with ``layout`` option. More information and parameters described in the `Embedding @@ -209,8 +258,9 @@ article `__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -245,8 +295,9 @@ Prepare image The data type of the image is float32 -Compile model and perform inerence -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Compile model and perform inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -256,25 +307,35 @@ Compile model and perform inerence result = compiled_model_mo_pp(mo_pp_input_tensor)[output_layer] -Setup preprocessing steps with Preprocessing API and perform inference ----------------------------------------------------------------------- +Setup preprocessing steps with Preprocessing API and perform inference. `⇑ <#top>`__ +############################################################################################################################### Intuitively, preprocessing API consists of the following parts: -- Tensor - declares user data format, like shape, layout, precision, color format from actual user’s data. -- Steps - describes sequence of preprocessing steps which need to be applied to user data. -- Model - specifies model data format. Usually, precision and shape are already known for model, only additional information, like layout can be specified. +- Tensor - declares user data format, like shape, layout, precision, + color format from actual user’s data. +- Steps - describes sequence of preprocessing steps which need to be + applied to user data. +- Model - specifies model data format. Usually, precision and shape are + already known for model, only additional information, like layout can + be specified. Graph modifications of a model shall be performed after the model is read from a drive and before it is loaded on the actual device. Pre-processing support following operations (please, see more details `here `__) -- Mean/Scale Normalization - Converting Precision - Converting layout -(transposing) - Resizing Image - Color Conversion - Custom Operations -Convert model to OpenVINO IR with Model Optimizer -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +- Mean/Scale Normalization +- Converting Precision +- Converting layout (transposing) +- Resizing Image +- Color Conversion +- Custom Operations + +Convert model to OpenVINO IR with model conversion API `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The options for preprocessing are not required. @@ -292,8 +353,9 @@ The options for preprocessing are not required. input_shape=[1,299,299,3]) serialize(ppp_model, str(ir_path)) -Create PrePostProcessor Object -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Create ``PrePostProcessor`` Object `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The `PrePostProcessor() `__ @@ -306,8 +368,9 @@ a model. ppp = PrePostProcessor(ppp_model) -Declare User’s Data Format -~~~~~~~~~~~~~~~~~~~~~~~~~~ +Declare User’s Data Format `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + To address particular input of a model/preprocessor, use the ``PrePostProcessor.input(input_name)`` method. If the model has only one @@ -324,9 +387,11 @@ for more information about parameters for overriding. Below is all the specified input information: -- Precision is ``U8``(unsigned 8-bit integer). -- Size is non-fixed, setup of one determined shape size can be done with ``.set_shape([1, 577, 800, 3])``. -- Layout is ``“NHWC”``. It means, for example: height=577, width=800, channels=3. +- Precision is ``U8`` (unsigned 8-bit integer). +- Size is non-fixed, setup of one determined shape size can be done + with ``.set_shape([1, 577, 800, 3])`` +- Layout is ``“NHWC”``. It means, for example: height=577, width=800, + channels=3. The height and width are necessary for resizing, and channels are needed for mean/scale normalization. @@ -345,12 +410,13 @@ for mean/scale normalization. .. parsed-literal:: - + -Declaring Model Layout -~~~~~~~~~~~~~~~~~~~~~~ +Declaring Model Layout `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Model input already has information about precision and shape. Preprocessing API is not intended to modify this. The only thing that @@ -374,12 +440,13 @@ may be specified is input data .. parsed-literal:: - + -Preprocessing Steps -~~~~~~~~~~~~~~~~~~~ +Preprocessing Steps `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Now, the sequence of preprocessing steps can be defined. For more information about preprocessing steps, see @@ -387,10 +454,14 @@ information about preprocessing steps, see Perform the following: -- Convert ``U8`` to ``FP32`` precision. -- Resize to height/width of a model. Be aware that if a model accepts dynamic size, for example, ``{?, 3, ?, ?}`` resize will not know how to resize the picture. Therefore, in this case, target height/ width should be specified. For more details, see also the `PreProcessSteps.resize() `__. -- Subtract mean from each channel. -- Divide each pixel data to appropriate scale value. +- Convert ``U8`` to ``FP32`` precision. +- Resize to height/width of a model. Be aware that if a model accepts + dynamic size, for example, ``{?, 3, ?, ?}`` resize will not know how + to resize the picture. Therefore, in this case, target height/ width + should be specified. For more details, see also the + `PreProcessSteps.resize() `__. +- Subtract mean from each channel. +- Divide each pixel data to appropriate scale value. There is no need to specify conversion layout. If layouts are different, then such conversion will be added explicitly. @@ -409,16 +480,17 @@ then such conversion will be added explicitly. .. parsed-literal:: - + -Integrating Steps into a Model -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Integrating Steps into a Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Once the preprocessing steps have been finished, the model can be -finally built. It is possible to display PrePostProcessor configuration -for debugging purposes. +finally built. It is possible to display ``PrePostProcessor`` +configuration for debugging purposes. .. code:: ipython3 @@ -439,8 +511,9 @@ for debugging purposes. -Load model and perform inference --------------------------------- +Load model and perform inference `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -457,19 +530,22 @@ Load model and perform inference ppp_input_tensor = prepare_image_api_preprocess(image_path) results = compiled_model_with_preprocess_api(ppp_input_tensor)[ppp_output_layer][0] -Fit image manually and perform inference ----------------------------------------- +Fit image manually and perform inference `⇑ <#top>`__ +############################################################################################################################### + + +Load the model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Load the model -~~~~~~~~~~~~~~ .. code:: ipython3 model = core.read_model(model=ir_path) compiled_model = core.compile_model(model=model, device_name=device.value) -Load image and fit it to model input -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Load image and fit it to model input `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -502,8 +578,9 @@ Load image and fit it to model input The data type of the image is float32 -Perform inference -~~~~~~~~~~~~~~~~~ +Perform inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -511,11 +588,13 @@ Perform inference result = compiled_model(input_tensor)[output_layer] -Compare results ---------------- +Compare results `⇑ <#top>`__ +############################################################################################################################### + + +Compare results on one image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Compare results on one image -~~~~~~~~~~~~~~~~~~~~~~~~~~~~ .. code:: ipython3 @@ -538,7 +617,7 @@ Compare results on one image imagenet_classes = ['background'] + imagenet_classes # get result for inference with preprocessing api - print("Result of inference for preprocessing with ModelOptimizer:") + print("Result of inference for preprocessing with Model Optimizer:") res = check_results(mo_pp_input_tensor, compiled_model_mo_pp, imagenet_classes) print("\n") @@ -556,7 +635,7 @@ Compare results on one image .. parsed-literal:: - Result of inference for preprocessing with ModelOptimizer: + Result of inference for preprocessing with Model Optimizer: n02099601 golden retriever, 0.56439 n02098413 Lhasa, Lhasa apso, 0.35731 n02108915 French bulldog, 0.00730 @@ -580,8 +659,9 @@ Compare results on one image n02100877 Irish setter, red setter, 0.00116 -Compare performance -~~~~~~~~~~~~~~~~~~~ +Compare performance `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -621,7 +701,7 @@ Compare performance .. parsed-literal:: - IR model in OpenVINO Runtime/CPU with preprocessing API: 0.0200 seconds per image, FPS: 49.90 - IR model in OpenVINO Runtime/CPU with preprocessing API: 0.0153 seconds per image, FPS: 65.52 - IR model in OpenVINO Runtime/CPU with preprocessing API: 0.0187 seconds per image, FPS: 53.59 + IR model in OpenVINO Runtime/CPU with preprocessing API: 0.0199 seconds per image, FPS: 50.13 + IR model in OpenVINO Runtime/CPU with preprocessing API: 0.0155 seconds per image, FPS: 64.58 + IR model in OpenVINO Runtime/CPU with preprocessing API: 0.0188 seconds per image, FPS: 53.27 diff --git a/docs/notebooks/118-optimize-preprocessing-with-output_files/118-optimize-preprocessing-with-output_12_1.png b/docs/notebooks/118-optimize-preprocessing-with-output_files/118-optimize-preprocessing-with-output_13_1.png similarity index 100% rename from docs/notebooks/118-optimize-preprocessing-with-output_files/118-optimize-preprocessing-with-output_12_1.png rename to docs/notebooks/118-optimize-preprocessing-with-output_files/118-optimize-preprocessing-with-output_13_1.png diff --git a/docs/notebooks/118-optimize-preprocessing-with-output_files/index.html b/docs/notebooks/118-optimize-preprocessing-with-output_files/index.html index 6e0c2d07dbd..8bc4294fea3 100644 --- a/docs/notebooks/118-optimize-preprocessing-with-output_files/index.html +++ b/docs/notebooks/118-optimize-preprocessing-with-output_files/index.html @@ -1,7 +1,7 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/118-optimize-preprocessing-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/118-optimize-preprocessing-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/118-optimize-preprocessing-with-output_files/


../
-118-optimize-preprocessing-with-output_12_1.png    12-Jul-2023 00:11              387941
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/118-optimize-preprocessing-with-output_files/


../
+118-optimize-preprocessing-with-output_13_1.png    16-Aug-2023 01:31              387941
 

diff --git a/docs/notebooks/119-tflite-to-openvino-with-output.rst b/docs/notebooks/119-tflite-to-openvino-with-output.rst index 7c2250bb48c..22209f4c1c7 100644 --- a/docs/notebooks/119-tflite-to-openvino-with-output.rst +++ b/docs/notebooks/119-tflite-to-openvino-with-output.rst @@ -1,25 +1,45 @@ Convert a Tensorflow Lite Model to OpenVINO™ ============================================ +.. _top: + `TensorFlow Lite `__, often referred to as TFLite, is an open source library developed for deploying machine learning models to edge devices. This short tutorial shows how to convert a TensorFlow Lite -`efficientnet-lite-b0 `__ +`EfficientNet-Lite-B0 `__ image classification model to OpenVINO `Intermediate Representation `__ (OpenVINO IR) format, using `Model Optimizer `__. After creating the OpenVINO IR, load the model in `OpenVINO -Runtime `__ -and do inference with a sample image. +Runtime `__ +and do inference with a sample image. -Preparation ------------ +**Table of contents**: + +- `Preparation <#preparation>`__ + + - `Install requirements <#install-requirements>`__ + - `Imports <#imports>`__ + +- `Download TFLite model <#download-tflite-model>`__ +- `Convert a Model to OpenVINO IR Format <#convert-a-model-to-openvino-ir-format>`__ +- `Load model using OpenVINO TensorFlow Lite Frontend <#load-model-using-openvino-tensorflow-lite-frontend>`__ +- `Run OpenVINO model inference <#run-openvino-model-inference>`__ + + - `Select inference device <#select-inference-device>`__ + +- `Estimate Model Performance <#estimate-model-performance>`__ + +Preparation `⇑ <#top>`__ +############################################################################################################################### + + +Install requirements `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Install requirements -~~~~~~~~~~~~~~~~~~~~ .. code:: ipython3 @@ -33,8 +53,9 @@ Install requirements filename='notebook_utils.py' ); -Imports -~~~~~~~ +Imports `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -46,8 +67,9 @@ Imports from notebook_utils import download_file, load_image -Download TFLite model ---------------------- +Download TFLite model `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -70,25 +92,26 @@ Download TFLite model .. parsed-literal:: - PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/119-tflite-to-openvino/model/efficientnet_lite0_fp32_2.tflite') + PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/119-tflite-to-openvino/model/efficientnet_lite0_fp32_2.tflite') -Convert a Model to OpenVINO IR Format -------------------------------------- +Convert a Model to OpenVINO IR Format `⇑ <#top>`__ +############################################################################################################################### -To convert the TFLite model to OpenVINO IR, OpenVINO Model Optimizer -Python API can be used. ``mo.convert_model`` function accept path to -TFLite model and returns OpenVINO Model class instance which represents -this model. Obtained model is ready to use and loading on device using -``compile_model`` or can be saved on disk using ``serialize`` function -reducing loading time for next running. Optionally, we can apply -compression to FP16 model weigths using ``compress_to_fp16=True`` option -and integrate preprocessing using this approach. See the `Model -Optimizer Developer -Guide `__ -for more information about Model Optimizer and TensorFlow Lite `models -suport `__. + +To convert the TFLite model to OpenVINO IR, model conversion Python API +can be used. ``mo.convert_model`` function accepts the path to the +TFLite model and returns an OpenVINO Model class instance which +represents this model. The obtained model is ready to use and to be +loaded on a device using ``compile_model`` or can be saved on a disk +using ``serialize`` function, reducing loading time for next running. +Optionally, we can apply compression to the FP16 model weights, using +the ``compress_to_fp16=True`` option and integrate preprocessing using +this approach. For more information about model conversion, see this +`page `__. +For TensorFlow Lite models support, refer to this +`tutorial `__. .. code:: ipython3 @@ -102,10 +125,11 @@ suport `__ +############################################################################################################################### -TensorFlow Lite models are supported via FrontEnd API. You may skip + +TensorFlow Lite models are supported via ``FrontEnd`` API. You may skip conversion to IR and read models directly by OpenVINO runtime API. For more examples supported formats reading via Frontend API, please look this `tutorial <../002-openvino-api>`__. @@ -116,8 +140,9 @@ this `tutorial <../002-openvino-api>`__. ov_model = core.read_model(tflite_model_path) -Run OpenVINO model inference ----------------------------- +Run OpenVINO model inference `⇑ <#top>`__ +############################################################################################################################### + We can find information about model input preprocessing in its `description `__ @@ -131,10 +156,11 @@ on `TensorFlow Hub `__. resized_image = image.resize((224, 224)) input_tensor = np.expand_dims((np.array(resized_image).astype(np.float32) - 127) / 128, 0) -Select inference device -~~~~~~~~~~~~~~~~~~~~~~~ +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -select device from dropdown list for running inference using OpenVINO + +Select device from dropdown list for running inference using OpenVINO: .. code:: ipython3 @@ -191,14 +217,14 @@ select device from dropdown list for running inference using OpenVINO Predicted label: n02109047 Great Dane with probability 0.715318 -Estimate Model Performance --------------------------- +Estimate Model Performance `⇑ <#top>`__ +############################################################################################################################### -`Benchmark -Tool `__ +`Benchmark Tool `__ is used to measure the inference performance of the model on CPU and GPU. + **NOTE**: For more accurate performance, it is recommended to run ``benchmark_app`` in a terminal/command prompt after closing other applications. Run ``benchmark_app -m model.xml -d CPU`` to benchmark @@ -233,7 +259,7 @@ GPU. [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.THROUGHPUT. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 21.86 ms + [ INFO ] Read model took 9.14 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] images (node: images) : f32 / [...] / [1,224,224,3] @@ -247,7 +273,7 @@ GPU. [ INFO ] Model outputs: [ INFO ] Softmax (node: 61) : f32 / [...] / [1,1000] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 183.81 ms + [ INFO ] Compile model took 151.57 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: [ INFO ] NETWORK_NAME: TensorFlow_Lite_Frontend_IR @@ -269,15 +295,15 @@ GPU. [ INFO ] Fill input 'images' with random values [Step 10/11] Measuring performance (Start inference asynchronously, 6 inference requests, limits: 15000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 7.41 ms + [ INFO ] First inference took 7.60 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 17364 iterations - [ INFO ] Duration: 15009.20 ms + [ INFO ] Count: 17526 iterations + [ INFO ] Duration: 15005.75 ms [ INFO ] Latency: - [ INFO ] Median: 5.04 ms - [ INFO ] Average: 5.05 ms - [ INFO ] Min: 3.22 ms - [ INFO ] Max: 15.36 ms - [ INFO ] Throughput: 1156.89 FPS + [ INFO ] Median: 5.00 ms + [ INFO ] Average: 5.00 ms + [ INFO ] Min: 3.28 ms + [ INFO ] Max: 14.83 ms + [ INFO ] Throughput: 1167.95 FPS diff --git a/docs/notebooks/119-tflite-to-openvino-with-output_files/index.html b/docs/notebooks/119-tflite-to-openvino-with-output_files/index.html index 8cf8feaae6c..902847c51f8 100644 --- a/docs/notebooks/119-tflite-to-openvino-with-output_files/index.html +++ b/docs/notebooks/119-tflite-to-openvino-with-output_files/index.html @@ -1,8 +1,8 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/119-tflite-to-openvino-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/119-tflite-to-openvino-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/119-tflite-to-openvino-with-output_files/


../
-119-tflite-to-openvino-with-output_16_1.jpg        12-Jul-2023 00:11               68170
-119-tflite-to-openvino-with-output_16_1.png        12-Jul-2023 00:11              621006
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/119-tflite-to-openvino-with-output_files/


../
+119-tflite-to-openvino-with-output_16_1.jpg        16-Aug-2023 01:31               68170
+119-tflite-to-openvino-with-output_16_1.png        16-Aug-2023 01:31              621006
 

diff --git a/docs/notebooks/120-tensorflow-object-detection-to-openvino-with-output.rst b/docs/notebooks/120-tensorflow-object-detection-to-openvino-with-output.rst index 1a5b7765b4d..a2f8edcf643 100644 --- a/docs/notebooks/120-tensorflow-object-detection-to-openvino-with-output.rst +++ b/docs/notebooks/120-tensorflow-object-detection-to-openvino-with-output.rst @@ -1,6 +1,8 @@ Convert a TensorFlow Object Detection Model to OpenVINO™ ======================================================== +.. _top: + `TensorFlow `__, or TF for short, is an open-source framework for machine learning. @@ -21,11 +23,33 @@ Representation `__. After creating the OpenVINO IR, load the model in `OpenVINO -Runtime `__ -and do inference with a sample image. +Runtime `__ +and do inference with a sample image. + +**Table of contents**: + +- `Prerequisites <#prerequisites>`__ +- `Imports <#imports>`__ +- `Settings <#settings>`__ +- `Download Model from TensorFlow Hub <#download-model-from-tensorflow-hub>`__ +- `Convert Model to OpenVINO IR <#convert-model-to-openvino-ir>`__ +- `Test Inference on the Converted Model <#test-inference-on-the-converted-model>`__ +- `Select inference device <#select-inference-device>`__ + + - `Load the Model <#load-the-model>`__ + - `Get Model Information <#get-model-information>`__ + - `Get an Image for Test Inference <#get-an-image-for-test-inference>`__ + - `Perform Inference <#perform-inference>`__ + - `Inference Result Visualization <#inference-result-visualization>`__ + +- `Next Steps <#next-steps>`__ + + - `Async inference pipeline <#async-inference-pipeline>`__ + - `Integration preprocessing to model <#integration-preprocessing-to-model>`__ + +Prerequisites `⇑ <#top>`__ +############################################################################################################################### -Prerequisites -------------- Install required packages: @@ -46,8 +70,9 @@ The notebook uses utility functions. The cell below will download the filename="notebook_utils.py", ); -Imports -------- +Imports `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -66,8 +91,9 @@ Imports from openvino.runtime import Core, serialize from openvino.tools import mo -Settings --------- +Settings `⇑ <#top>`__ +############################################################################################################################### + Define model related variables and create corresponding directories: @@ -93,8 +119,9 @@ Define model related variables and create corresponding directories: tf_model_archive_filename = f"{model_name}.tar.gz" -Download Model from TensorFlow Hub ----------------------------------- +Download Model from TensorFlow Hub `⇑ <#top>`__ +############################################################################################################################### + Download archive with TensorFlow Object Detection model (`faster_rcnn_resnet50_v1_640x640 `__) @@ -119,7 +146,7 @@ from TensorFlow Hub: .. parsed-literal:: - PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/120-tensorflow-object-detection-to-openvino/model/tf/faster_rcnn_resnet50_v1_640x640.tar.gz') + PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/120-tensorflow-object-detection-to-openvino/model/tf/faster_rcnn_resnet50_v1_640x640.tar.gz') @@ -132,8 +159,9 @@ Extract TensorFlow Object Detection model from the downloaded archive: with tarfile.open(tf_model_dir / tf_model_archive_filename) as file: file.extractall(path=tf_model_dir) -Convert Model to OpenVINO IR ----------------------------- +Convert Model to OpenVINO IR `⇑ <#top>`__ +############################################################################################################################### + OpenVINO Model Optimizer Python API can be used to convert the TensorFlow model to OpenVINO IR. @@ -143,7 +171,7 @@ returns OpenVINO Model class instance which represents this model. Also we need to provide model input shape (``input_shape``) that is described at `model overview page on TensorFlow Hub `__. -Optionally, we can apply compression to FP16 model weigths using +Optionally, we can apply compression to FP16 model weights using ``compress_to_fp16=True`` option and integrate preprocessing using this approach. @@ -154,7 +182,7 @@ when the model is run in the future. See the `Model Optimizer Developer Guide `__ for more information about Model Optimizer and TensorFlow `models -suport `__. +support `__. .. code:: ipython3 @@ -166,13 +194,15 @@ suport `__ +############################################################################################################################### -Select inference device ------------------------ -select device from dropdown list for running inference using OpenVINO +Select inference device `⇑ <#top>`__ +############################################################################################################################### + + +Select device from dropdown list for running inference using OpenVINO: .. code:: ipython3 @@ -197,8 +227,9 @@ select device from dropdown list for running inference using OpenVINO -Load the Model -~~~~~~~~~~~~~~ +Load the Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -206,8 +237,9 @@ Load the Model openvino_ir_model = core.read_model(openvino_ir_path) compiled_model = core.compile_model(model=openvino_ir_model, device_name=device.value) -Get Model Information -~~~~~~~~~~~~~~~~~~~~~ +Get Model Information `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Faster R-CNN with Resnet-50 V1 object detection model has one input - a three-channel image of variable size. The input tensor shape is @@ -215,14 +247,22 @@ three-channel image of variable size. The input tensor shape is Model output dictionary contains several tensors: -- ``num_detections`` - the number of detections in ``[N]`` format. -- ``detection_boxes`` - bounding box coordinates for all ``N`` detections in ``[ymin, xmin, ymax, xmax]`` format. -- ``detection_classes`` - ``N`` detection class indexes size from the label file. -- ``detection_scores``- ``N`` detection scores (confidence) for each detected class. -- ``raw_detection_boxes`` - decoded detection boxes without Non-Max suppression. -- ``raw_detection_scores`` - class score logits for raw detection boxes. -- ``detection_anchor_indices`` - the anchor indices of the detections after NMS. -- ``detection_multiclass_scores`` - class score distribution (including background) for detection boxes in the image including background class. +- ``num_detections`` - the number of detections in ``[N]`` format. +- ``detection_boxes`` - bounding box coordinates for all ``N`` + detections in ``[ymin, xmin, ymax, xmax]`` format. +- ``detection_classes`` - ``N`` detection class indexes size from the + label file. +- ``detection_scores`` - ``N`` detection scores (confidence) for each + detected class. +- ``raw_detection_boxes`` - decoded detection boxes without Non-Max + suppression. +- ``raw_detection_scores`` - class score logits for raw detection + boxes. +- ``detection_anchor_indices`` - the anchor indices of the detections + after NMS. +- ``detection_multiclass_scores`` - class score distribution (including + background) for detection boxes in the image including background + class. In this tutorial we will mostly use ``detection_boxes``, ``detection_classes``, ``detection_scores`` tensors. It is important to @@ -266,8 +306,9 @@ for more information about model inputs, outputs and their formats. -Get an Image for Test Inference -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Get an Image for Test Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Load and save an image: @@ -292,7 +333,7 @@ Load and save an image: .. parsed-literal:: - PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/120-tensorflow-object-detection-to-openvino/data/coco_bike.jpg') + PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/120-tensorflow-object-detection-to-openvino/data/coco_bike.jpg') @@ -320,7 +361,7 @@ Read the image, resize and convert it to the input shape of the network: .. parsed-literal:: - + @@ -328,8 +369,9 @@ Read the image, resize and convert it to the input shape of the network: .. image:: 120-tensorflow-object-detection-to-openvino-with-output_files/120-tensorflow-object-detection-to-openvino-with-output_25_1.png -Perform Inference -~~~~~~~~~~~~~~~~~ +Perform Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -439,8 +481,9 @@ outputs will be used. image_detections_num: [300.] -Inference Result Visualization -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Inference Result Visualization `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Define utility functions to visualize the inference results @@ -580,7 +623,7 @@ Zoo `__: .. parsed-literal:: - PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/120-tensorflow-object-detection-to-openvino/data/coco_91cl.txt') + PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/120-tensorflow-object-detection-to-openvino/data/coco_91cl.txt') @@ -619,24 +662,25 @@ original test image: .. image:: 120-tensorflow-object-detection-to-openvino-with-output_files/120-tensorflow-object-detection-to-openvino-with-output_38_0.png -Next Steps ----------- +Next Steps `⇑ <#top>`__ +############################################################################################################################### + This section contains suggestions on how to additionally improve the performance of your application using OpenVINO. -Async inference pipeline -~~~~~~~~~~~~~~~~~~~~~~~~ +Async inference pipeline `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -The key advantage of the Async API is that when a device is busy with -inference, the application can perform other tasks in parallel (for -example, populating inputs or scheduling other requests) rather than -wait for the current inference to complete first. To understand how to -perform async inference using openvino, refer to the `Async API -tutorial <115-async-api-with-output.html>`__. +The key advantage of the Async API is that when a device is busy with inference, +the application can perform other tasks in parallel (for example, populating inputs or +scheduling other requests) rather than wait for the current inference to +complete first. To understand how to perform async inference using +openvino, refer to the `Async API tutorial <115-async-api-with-output.html>`__. + +Integration preprocessing to model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Integration preprocessing to model -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Preprocessing API enables making preprocessing a part of the model reducing application code and dependency on additional image processing diff --git a/docs/notebooks/120-tensorflow-object-detection-to-openvino-with-output_files/120-tensorflow-object-detection-to-openvino-with-output_38_0.png b/docs/notebooks/120-tensorflow-object-detection-to-openvino-with-output_files/120-tensorflow-object-detection-to-openvino-with-output_38_0.png index 2be72c724a5..dff0d72c697 100644 --- a/docs/notebooks/120-tensorflow-object-detection-to-openvino-with-output_files/120-tensorflow-object-detection-to-openvino-with-output_38_0.png +++ b/docs/notebooks/120-tensorflow-object-detection-to-openvino-with-output_files/120-tensorflow-object-detection-to-openvino-with-output_38_0.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:59da4a179a69c8527e526a4b705f9d2ac5225cb90e2489353e36340dd12b481c -size 391541 +oid sha256:b7874b5d950db3016c015449880bbc25d55dcbb66b14f1fd54593352f044b3a2 +size 391330 diff --git a/docs/notebooks/120-tensorflow-object-detection-to-openvino-with-output_files/index.html b/docs/notebooks/120-tensorflow-object-detection-to-openvino-with-output_files/index.html index ab7a24a2eef..e98e97f8542 100644 --- a/docs/notebooks/120-tensorflow-object-detection-to-openvino-with-output_files/index.html +++ b/docs/notebooks/120-tensorflow-object-detection-to-openvino-with-output_files/index.html @@ -1,8 +1,8 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/120-tensorflow-object-detection-to-openvino-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/120-tensorflow-object-detection-to-openvino-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/120-tensorflow-object-detection-to-openvino-with-output_files/


../
-120-tensorflow-object-detection-to-openvino-wit..> 12-Jul-2023 00:11              395346
-120-tensorflow-object-detection-to-openvino-wit..> 12-Jul-2023 00:11              391541
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/120-tensorflow-object-detection-to-openvino-with-output_files/


../
+120-tensorflow-object-detection-to-openvino-wit..> 16-Aug-2023 01:31              395346
+120-tensorflow-object-detection-to-openvino-wit..> 16-Aug-2023 01:31              391330
 

diff --git a/docs/notebooks/121-convert-to-openvino-with-output.rst b/docs/notebooks/121-convert-to-openvino-with-output.rst new file mode 100644 index 00000000000..5da2d317e3a --- /dev/null +++ b/docs/notebooks/121-convert-to-openvino-with-output.rst @@ -0,0 +1,1474 @@ +OpenVINO™ model conversion API +============================== + +This notebook shows how to convert a model from original framework +format to OpenVINO Intermediate Representation (IR). + +**Table of contents**: + +- `OpenVINO IR format <#openvino-ir-format>`__ +- `IR preparation with Python conversion API and Model Optimizer + command-line + tool <#ir-preparation-with-python-conversion-api-and-model-optimizer-command-line-tool>`__ +- `Fetching example models <#fetching-example-models>`__ +- `Basic conversion <#basic-conversion>`__ +- `Model conversion parameters <#model-conversion-parameters>`__ + + - `Setting Input Shapes <#setting-input-shapes>`__ + - `Cutting Off Parts of a Model <#cutting-off-parts-of-a-model>`__ + - `Embedding Preprocessing + Computation <#embedding-preprocessing-computation>`__ + + - `Specifying Layout <#specifying-layout>`__ + - `Changing Model Layout <#changing-model-layout>`__ + - `Specifying Mean and Scale + Values <#specifying-mean-and-scale-values>`__ + - `Reversing Input Channels <#reversing-input-channels>`__ + + - `Compressing a Model to FP16 <#compressing-a-model-to-fp16>`__ + +- `Convert Models Represented as Python + Objects <#convert-models-represented-as-python-objects>`__ + +.. code-block:: ipython3 + :force: + + # Required imports. Please execute this cell first. + ! pip install -q --find-links https://download.pytorch.org/whl/torch_stable.html \ + "openvino-dev>=2023.0.1" \ + "requests" \ + "tqdm" \ + "transformers[onnx]>=4.21.1" \ + "torch==1.13.1; sys_platform == 'darwin'" \ + "torch==1.13.1+cpu; sys_platform == 'linux' or platform_system == 'Windows'" \ + "torchvision==0.14.1; sys_platform == 'darwin'" \ + "torchvision==0.14.1+cpu; sys_platform == 'linux' or platform_system == 'Windows'" + +OpenVINO IR format +------------------ + +OpenVINO `Intermediate Representation +(IR) `__ is the +proprietary model format of OpenVINO. It is produced after converting a +model with model conversion API. Model conversion API translates the +frequently used deep learning operations to their respective similar +representation in OpenVINO and tunes them with the associated weights +and biases from the trained model. The resulting IR contains two files: +an ``.xml`` file, containing information about network topology, and a +``.bin`` file, containing the weights and biases binary data. + +IR preparation with Python conversion API and Model Optimizer command-line tool +------------------------------------------------------------------------------- + +There are two ways to convert a model from the original framework format +to OpenVINO IR: Python conversion API and Model Optimizer command-line +tool. You can choose one of them based on whichever is most convenient +for you. There should not be any differences in the results of model +conversion if the same set of parameters is used. For more details, +refer to `Model +Preparation `__ +documentation. + +.. code:: ipython3 + + # Model Optimizer CLI tool parameters description + + ! mo --help + + +.. parsed-literal:: + + usage: main.py [options] + + optional arguments: + -h, --help show this help message and exit + --framework FRAMEWORK + Name of the framework used to train the input model. + + Framework-agnostic parameters: + --model_name MODEL_NAME, -n MODEL_NAME + Model_name parameter passed to the final create_ir + transform. This parameter is used to name a network in + a generated IR and output .xml/.bin files. + --output_dir OUTPUT_DIR, -o OUTPUT_DIR + Directory that stores the generated IR. By default, it + is the directory from where the Model Optimizer is + launched. + --freeze_placeholder_with_value FREEZE_PLACEHOLDER_WITH_VALUE + Replaces input layer with constant node with provided + value, for example: "node_name->True". It will be + DEPRECATED in future releases. Use --input option to + specify a value for freezing. + --static_shape Enables IR generation for fixed input shape (folding + `ShapeOf` operations and shape-calculating sub-graphs + to `Constant`). Changing model input shape using the + OpenVINO Runtime API in runtime may fail for such an + IR. + --use_new_frontend Force the usage of new Frontend of Model Optimizer for + model conversion into IR. The new Frontend is C++ + based and is available for ONNX* and PaddlePaddle* + models. Model optimizer uses new Frontend for ONNX* + and PaddlePaddle* by default that means + `--use_new_frontend` and `--use_legacy_frontend` + options are not specified. + --use_legacy_frontend + Force the usage of legacy Frontend of Model Optimizer + for model conversion into IR. The legacy Frontend is + Python based and is available for TensorFlow*, ONNX*, + MXNet*, Caffe*, and Kaldi* models. + --input_model INPUT_MODEL, -w INPUT_MODEL, -m INPUT_MODEL + Tensorflow*: a file with a pre-trained model (binary + or text .pb file after freezing). Caffe*: a model + proto file with model weights. + --input INPUT Quoted list of comma-separated input nodes names with + shapes, data types, and values for freezing. The order + of inputs in converted model is the same as order of + specified operation names. The shape and value are + specified as comma-separated lists. The data type of + input node is specified in braces and can have one of + the values: f64 (float64), f32 (float32), f16 + (float16), i64 (int64), i32 (int32), u8 (uint8), + boolean (bool). Data type is optional. If it's not + specified explicitly then there are two options: if + input node is a parameter, data type is taken from the + original node dtype, if input node is not a parameter, + data type is set to f32. Example, to set `input_1` + with shape [1,100], and Parameter node `sequence_len` + with scalar input with value `150`, and boolean input + `is_training` with `False` value use the following + format: + "input_1[1,100],sequence_len->150,is_training->False". + Another example, use the following format to set input + port 0 of the node `node_name1` with the shape [3,4] + as an input node and freeze output port 1 of the node + "node_name2" with the value [20,15] of the int32 type + and shape [2]: + "0:node_name1[3,4],node_name2:1[2]{i32}->[20,15]". + --output OUTPUT The name of the output operation of the model or list + of names. For TensorFlow*, do not add :0 to this + name.The order of outputs in converted model is the + same as order of specified operation names. + --input_shape INPUT_SHAPE + Input shape(s) that should be fed to an input node(s) + of the model. Shape is defined as a comma-separated + list of integer numbers enclosed in parentheses or + square brackets, for example [1,3,227,227] or + (1,227,227,3), where the order of dimensions depends + on the framework input layout of the model. For + example, [N,C,H,W] is used for ONNX* models and + [N,H,W,C] for TensorFlow* models. The shape can + contain undefined dimensions (? or -1) and should fit + the dimensions defined in the input operation of the + graph. Boundaries of undefined dimension can be + specified with ellipsis, for example + [1,1..10,128,128]. One boundary can be undefined, for + example [1,..100] or [1,3,1..,1..]. If there are + multiple inputs in the model, --input_shape should + contain definition of shape for each input separated + by a comma, for example: [1,3,227,227],[2,4] for a + model with two inputs with 4D and 2D shapes. + Alternatively, specify shapes with the --input option. + --batch BATCH, -b BATCH + Set batch size. It applies to 1D or higher dimension + inputs. The default dimension index for the batch is + zero. Use a label 'n' in --layout or --source_layout + option to set the batch dimension. For example, + "x(hwnc)" defines the third dimension to be the batch. + --mean_values MEAN_VALUES + Mean values to be used for the input image per + channel. Values to be provided in the (R,G,B) or + [R,G,B] format. Can be defined for desired input of + the model, for example: "--mean_values + data[255,255,255],info[255,255,255]". The exact + meaning and order of channels depend on how the + original model was trained. + --scale_values SCALE_VALUES + Scale values to be used for the input image per + channel. Values are provided in the (R,G,B) or [R,G,B] + format. Can be defined for desired input of the model, + for example: "--scale_values + data[255,255,255],info[255,255,255]". The exact + meaning and order of channels depend on how the + original model was trained. If both --mean_values and + --scale_values are specified, the mean is subtracted + first and then scale is applied regardless of the + order of options in command line. + --scale SCALE, -s SCALE + All input values coming from original network inputs + will be divided by this value. When a list of inputs + is overridden by the --input parameter, this scale is + not applied for any input that does not match with the + original input of the model. If both --mean_values and + --scale are specified, the mean is subtracted first + and then scale is applied regardless of the order of + options in command line. + --reverse_input_channels [REVERSE_INPUT_CHANNELS] + Switch the input channels order from RGB to BGR (or + vice versa). Applied to original inputs of the model + if and only if a number of channels equals 3. When + --mean_values/--scale_values are also specified, + reversing of channels will be applied to user's input + data first, so that numbers in --mean_values and + --scale_values go in the order of channels used in the + original model. In other words, if both options are + specified, then the data flow in the model looks as + following: Parameter -> ReverseInputChannels -> Mean + apply-> Scale apply -> the original body of the model. + --source_layout SOURCE_LAYOUT + Layout of the input or output of the model in the + framework. Layout can be specified in the short form, + e.g. nhwc, or in complex form, e.g. "[n,h,w,c]". + Example for many names: "in_name1([n,h,w,c]),in_name2( + nc),out_name1(n),out_name2(nc)". Layout can be + partially defined, "?" can be used to specify + undefined layout for one dimension, "..." can be used + to specify undefined layout for multiple dimensions, + for example "?c??", "nc...", "n...c", etc. + --target_layout TARGET_LAYOUT + Same as --source_layout, but specifies target layout + that will be in the model after processing by + ModelOptimizer. + --layout LAYOUT Combination of --source_layout and --target_layout. + Can't be used with either of them. If model has one + input it is sufficient to specify layout of this + input, for example --layout nhwc. To specify layouts + of many tensors, names must be provided, for example: + --layout "name1(nchw),name2(nc)". It is possible to + instruct ModelOptimizer to change layout, for example: + --layout "name1(nhwc->nchw),name2(cn->nc)". Also "*" + in long layout form can be used to fuse dimensions, + for example "[n,c,...]->[n*c,...]". + --compress_to_fp16 [COMPRESS_TO_FP16] + If the original model has FP32 weights or biases, they + are compressed to FP16. All intermediate data is kept + in original precision. Option can be specified alone + as "--compress_to_fp16", or explicit True/False values + can be set, for example: "--compress_to_fp16=False", + or "--compress_to_fp16=True" + --extensions EXTENSIONS + Paths or a comma-separated list of paths to libraries + (.so or .dll) with extensions. For the legacy MO path + (if `--use_legacy_frontend` is used), a directory or a + comma-separated list of directories with extensions + are supported. To disable all extensions including + those that are placed at the default location, pass an + empty string. + --transform TRANSFORM + Apply additional transformations. Usage: "--transform + transformation_name1[args],transformation_name2..." + where [args] is key=value pairs separated by + semicolon. Examples: "--transform LowLatency2" or "-- + transform Pruning" or "--transform + LowLatency2[use_const_initializer=False]" or "-- + transform "MakeStateful[param_res_names= {'input_name_ + 1':'output_name_1','input_name_2':'output_name_2'}]" + Available transformations: "LowLatency2", + "MakeStateful", "Pruning" + --transformations_config TRANSFORMATIONS_CONFIG + Use the configuration file with transformations + description. Transformations file can be specified as + relative path from the current directory, as absolute + path or as arelative path from the mo root directory. + --silent [SILENT] Prevent any output messages except those that + correspond to log level equals ERROR, that can be set + with the following option: --log_level. By default, + log level is already ERROR. + --log_level {CRITICAL,ERROR,WARN,WARNING,INFO,DEBUG,NOTSET} + Logger level of logging massages from MO. Expected one + of ['CRITICAL', 'ERROR', 'WARN', 'WARNING', 'INFO', + 'DEBUG', 'NOTSET']. + --version Version of Model Optimizer + --progress [PROGRESS] + Enable model conversion progress display. + --stream_output [STREAM_OUTPUT] + Switch model conversion progress display to a + multiline mode. + + TensorFlow*-specific parameters: + --input_model_is_text [INPUT_MODEL_IS_TEXT] + TensorFlow*: treat the input model file as a text + protobuf format. If not specified, the Model Optimizer + treats it as a binary file by default. + --input_checkpoint INPUT_CHECKPOINT + TensorFlow*: variables file to load. + --input_meta_graph INPUT_META_GRAPH + Tensorflow*: a file with a meta-graph of the model + before freezing + --saved_model_dir SAVED_MODEL_DIR + TensorFlow*: directory with a model in SavedModel + format of TensorFlow 1.x or 2.x version. + --saved_model_tags SAVED_MODEL_TAGS + Group of tag(s) of the MetaGraphDef to load, in string + format, separated by ','. For tag-set contains + multiple tags, all tags must be passed in. + --tensorflow_custom_operations_config_update TENSORFLOW_CUSTOM_OPERATIONS_CONFIG_UPDATE + TensorFlow*: update the configuration file with node + name patterns with input/output nodes information. + --tensorflow_object_detection_api_pipeline_config TENSORFLOW_OBJECT_DETECTION_API_PIPELINE_CONFIG + TensorFlow*: path to the pipeline configuration file + used to generate model created with help of Object + Detection API. + --tensorboard_logdir TENSORBOARD_LOGDIR + TensorFlow*: dump the input graph to a given directory + that should be used with TensorBoard. + --tensorflow_custom_layer_libraries TENSORFLOW_CUSTOM_LAYER_LIBRARIES + TensorFlow*: comma separated list of shared libraries + with TensorFlow* custom operations implementation. + + Caffe*-specific parameters: + --input_proto INPUT_PROTO, -d INPUT_PROTO + Deploy-ready prototxt file that contains a topology + structure and layer attributes + --caffe_parser_path CAFFE_PARSER_PATH + Path to Python Caffe* parser generated from + caffe.proto + --k K Path to CustomLayersMapping.xml to register custom + layers + --disable_omitting_optional [DISABLE_OMITTING_OPTIONAL] + Disable omitting optional attributes to be used for + custom layers. Use this option if you want to transfer + all attributes of a custom layer to IR. Default + behavior is to transfer the attributes with default + values and the attributes defined by the user to IR. + --enable_flattening_nested_params [ENABLE_FLATTENING_NESTED_PARAMS] + Enable flattening optional params to be used for + custom layers. Use this option if you want to transfer + attributes of a custom layer to IR with flattened + nested parameters. Default behavior is to transfer the + attributes without flattening nested parameters. + + MXNet-specific parameters: + --input_symbol INPUT_SYMBOL + Symbol file (for example, model-symbol.json) that + contains a topology structure and layer attributes + --nd_prefix_name ND_PREFIX_NAME + Prefix name for args.nd and argx.nd files. + --pretrained_model_name PRETRAINED_MODEL_NAME + Name of a pretrained MXNet model without extension and + epoch number. This model will be merged with args.nd + and argx.nd files + --save_params_from_nd [SAVE_PARAMS_FROM_ND] + Enable saving built parameters file from .nd files + --legacy_mxnet_model [LEGACY_MXNET_MODEL] + Enable MXNet loader to make a model compatible with + the latest MXNet version. Use only if your model was + trained with MXNet version lower than 1.0.0 + --enable_ssd_gluoncv [ENABLE_SSD_GLUONCV] + Enable pattern matchers replacers for converting + gluoncv ssd topologies. + + Kaldi-specific parameters: + --counts COUNTS Path to the counts file + --remove_output_softmax [REMOVE_OUTPUT_SOFTMAX] + Removes the SoftMax layer that is the output layer + --remove_memory [REMOVE_MEMORY] + Removes the Memory layer and use additional inputs + outputs instead + + +.. code:: ipython3 + + # Python conversion API parameters description + from openvino.tools import mo + + + mo.convert_model(help=True) + + +.. parsed-literal:: + + Optional parameters: + --help + Print available parameters. + --framework + Name of the framework used to train the input model. + + Framework-agnostic parameters: + --input_model + Model object in original framework (PyTorch, Tensorflow) or path to + model file. + Tensorflow*: a file with a pre-trained model (binary or text .pb file + after freezing). + Caffe*: a model proto file with model weights + + Supported formats of input model: + + PyTorch + torch.nn.Module + torch.jit.ScriptModule + torch.jit.ScriptFunction + + TF + tf.compat.v1.Graph + tf.compat.v1.GraphDef + tf.compat.v1.wrap_function + tf.compat.v1.session + + TF2 / Keras + tf.keras.Model + tf.keras.layers.Layer + tf.function + tf.Module + tf.train.checkpoint + --input + Input can be set by passing a list of InputCutInfo objects or by a list + of tuples. Each tuple can contain optionally input name, input + type or input shape. Example: input=("op_name", PartialShape([-1, + 3, 100, 100]), Type(np.float32)). Alternatively input can be set by + a string or list of strings of the following format. Quoted list of comma-separated + input nodes names with shapes, data types, and values for freezing. + If operation names are specified, the order of inputs in converted + model will be the same as order of specified operation names (applicable + for TF2, ONNX, MxNet). + The shape and value are specified as comma-separated lists. The data + type of input node is specified + in braces and can have one of the values: f64 (float64), f32 (float32), + f16 (float16), i64 + (int64), i32 (int32), u8 (uint8), boolean (bool). Data type is optional. + If it's not specified explicitly then there are two options: if input + node is a parameter, data type is taken from the original node dtype, + if input node is not a parameter, data type is set to f32. Example, to set + `input_1` with shape [1,100], and Parameter node `sequence_len` with + scalar input with value `150`, and boolean input `is_training` with + `False` value use the following format: "input_1[1,100],sequence_len->150,is_training->False". + Another example, use the following format to set input port 0 of the node + `node_name1` with the shape [3,4] as an input node and freeze output + port 1 of the node `node_name2` with the value [20,15] of the int32 type + and shape [2]: "0:node_name1[3,4],node_name2:1[2]{i32}->[20,15]". + + --output + The name of the output operation of the model or list of names. For TensorFlow*, + do not add :0 to this name.The order of outputs in converted model is the + same as order of specified operation names. + --input_shape + Input shape(s) that should be fed to an input node(s) of the model. Input + shapes can be defined by passing a list of objects of type PartialShape, + Shape, [Dimension, ...] or [int, ...] or by a string of the following + format. Shape is defined as a comma-separated list of integer numbers + enclosed in parentheses or square brackets, for example [1,3,227,227] + or (1,227,227,3), where the order of dimensions depends on the framework + input layout of the model. For example, [N,C,H,W] is used for ONNX* models + and [N,H,W,C] for TensorFlow* models. The shape can contain undefined + dimensions (? or -1) and should fit the dimensions defined in the input + operation of the graph. Boundaries of undefined dimension can be specified + with ellipsis, for example [1,1..10,128,128]. One boundary can be + undefined, for example [1,..100] or [1,3,1..,1..]. If there are multiple + inputs in the model, --input_shape should contain definition of shape + for each input separated by a comma, for example: [1,3,227,227],[2,4] + for a model with two inputs with 4D and 2D shapes. Alternatively, specify + shapes with the --input option. + --batch + Set batch size. It applies to 1D or higher dimension inputs. + The default dimension index for the batch is zero. + Use a label 'n' in --layout or --source_layout option to set the batch + dimension. + For example, "x(hwnc)" defines the third dimension to be the batch. + + --mean_values + Mean values to be used for the input image per channel. Mean values can + be set by passing a dictionary, where key is input name and value is mean + value. For example mean_values={'data':[255,255,255],'info':[255,255,255]}. + Or mean values can be set by a string of the following format. Values to + be provided in the (R,G,B) or [R,G,B] format. Can be defined for desired + input of the model, for example: "--mean_values data[255,255,255],info[255,255,255]". + The exact meaning and order of channels depend on how the original model + was trained. + --scale_values + Scale values to be used for the input image per channel. Scale values + can be set by passing a dictionary, where key is input name and value is + scale value. For example scale_values={'data':[255,255,255],'info':[255,255,255]}. + Or scale values can be set by a string of the following format. Values + are provided in the (R,G,B) or [R,G,B] format. Can be defined for desired + input of the model, for example: "--scale_values data[255,255,255],info[255,255,255]". + The exact meaning and order of channels depend on how the original model + was trained. If both --mean_values and --scale_values are specified, + the mean is subtracted first and then scale is applied regardless of + the order of options in command line. + --scale + All input values coming from original network inputs will be divided + by this value. When a list of inputs is overridden by the --input parameter, + this scale is not applied for any input that does not match with the original + input of the model. If both --mean_values and --scale are specified, + the mean is subtracted first and then scale is applied regardless of + the order of options in command line. + --reverse_input_channels + Switch the input channels order from RGB to BGR (or vice versa). Applied + to original inputs of the model if and only if a number of channels equals + 3. When --mean_values/--scale_values are also specified, reversing + of channels will be applied to user's input data first, so that numbers + in --mean_values and --scale_values go in the order of channels used + in the original model. In other words, if both options are specified, + then the data flow in the model looks as following: Parameter -> ReverseInputChannels + -> Mean apply-> Scale apply -> the original body of the model. + --source_layout + Layout of the input or output of the model in the framework. Layout can + be set by passing a dictionary, where key is input name and value is LayoutMap + object. Or layout can be set by string of the following format. Layout + can be specified in the short form, e.g. nhwc, or in complex form, e.g. + "[n,h,w,c]". Example for many names: "in_name1([n,h,w,c]),in_name2(nc),out_name1(n),out_name2(nc)". + Layout can be partially defined, "?" can be used to specify undefined + layout for one dimension, "..." can be used to specify undefined layout + for multiple dimensions, for example "?c??", "nc...", "n...c", etc. + + --target_layout + Same as --source_layout, but specifies target layout that will be in + the model after processing by ModelOptimizer. + --layout + Combination of --source_layout and --target_layout. Can't be used + with either of them. If model has one input it is sufficient to specify + layout of this input, for example --layout nhwc. To specify layouts + of many tensors, names must be provided, for example: --layout "name1(nchw),name2(nc)". + It is possible to instruct ModelOptimizer to change layout, for example: + --layout "name1(nhwc->nchw),name2(cn->nc)". + Also "*" in long layout form can be used to fuse dimensions, for example + "[n,c,...]->[n*c,...]". + --compress_to_fp16 + If the original model has FP32 weights or biases, they are compressed + to FP16. All intermediate data is kept in original precision. Option + can be specified alone as "--compress_to_fp16", or explicit True/False + values can be set, for example: "--compress_to_fp16=False", or "--compress_to_fp16=True" + + --extensions + Paths to libraries (.so or .dll) with extensions, comma-separated + list of paths, objects derived from BaseExtension class or lists of + objects. For the legacy MO path (if `--use_legacy_frontend` is used), + a directory or a comma-separated list of directories with extensions + are supported. To disable all extensions including those that are placed + at the default location, pass an empty string. + --transform + Apply additional transformations. 'transform' can be set by a list + of tuples, where the first element is transform name and the second element + is transform parameters. For example: [('LowLatency2', {{'use_const_initializer': + False}}), ...]"--transform transformation_name1[args],transformation_name2..." + where [args] is key=value pairs separated by semicolon. Examples: + "--transform LowLatency2" or + "--transform Pruning" or + "--transform LowLatency2[use_const_initializer=False]" or + "--transform "MakeStateful[param_res_names= + {'input_name_1':'output_name_1','input_name_2':'output_name_2'}]"" + Available transformations: "LowLatency2", "MakeStateful", "Pruning" + + --transformations_config + Use the configuration file with transformations description or pass + object derived from BaseExtension class. Transformations file can + be specified as relative path from the current directory, as absolute + path or as relative path from the mo root directory. + --silent + Prevent any output messages except those that correspond to log level + equals ERROR, that can be set with the following option: --log_level. + By default, log level is already ERROR. + --log_level + Logger level of logging massages from MO. + Expected one of ['CRITICAL', 'ERROR', 'WARN', 'WARNING', 'INFO', + 'DEBUG', 'NOTSET']. + --version + Version of Model Optimizer + --progress + Enable model conversion progress display. + --stream_output + Switch model conversion progress display to a multiline mode. + + PyTorch-specific parameters: + --example_input + Sample of model input in original framework. For PyTorch it can be torch.Tensor. + + + TensorFlow*-specific parameters: + --input_model_is_text + TensorFlow*: treat the input model file as a text protobuf format. If + not specified, the Model Optimizer treats it as a binary file by default. + + --input_checkpoint + TensorFlow*: variables file to load. + --input_meta_graph + Tensorflow*: a file with a meta-graph of the model before freezing + --saved_model_dir + TensorFlow*: directory with a model in SavedModel format of TensorFlow + 1.x or 2.x version. + --saved_model_tags + Group of tag(s) of the MetaGraphDef to load, in string format, separated + by ','. For tag-set contains multiple tags, all tags must be passed in. + + --tensorflow_custom_operations_config_update + TensorFlow*: update the configuration file with node name patterns + with input/output nodes information. + --tensorflow_object_detection_api_pipeline_config + TensorFlow*: path to the pipeline configuration file used to generate + model created with help of Object Detection API. + --tensorboard_logdir + TensorFlow*: dump the input graph to a given directory that should be + used with TensorBoard. + --tensorflow_custom_layer_libraries + TensorFlow*: comma separated list of shared libraries with TensorFlow* + custom operations implementation. + + MXNet-specific parameters: + --input_symbol + Symbol file (for example, model-symbol.json) that contains a topology + structure and layer attributes + --nd_prefix_name + Prefix name for args.nd and argx.nd files. + --pretrained_model_name + Name of a pretrained MXNet model without extension and epoch number. + This model will be merged with args.nd and argx.nd files + --save_params_from_nd + Enable saving built parameters file from .nd files + --legacy_mxnet_model + Enable MXNet loader to make a model compatible with the latest MXNet + version. Use only if your model was trained with MXNet version lower + than 1.0.0 + --enable_ssd_gluoncv + Enable pattern matchers replacers for converting gluoncv ssd topologies. + + + Caffe*-specific parameters: + --input_proto + Deploy-ready prototxt file that contains a topology structure and + layer attributes + --caffe_parser_path + Path to Python Caffe* parser generated from caffe.proto + --k + Path to CustomLayersMapping.xml to register custom layers + --disable_omitting_optional + Disable omitting optional attributes to be used for custom layers. + Use this option if you want to transfer all attributes of a custom layer + to IR. Default behavior is to transfer the attributes with default values + and the attributes defined by the user to IR. + --enable_flattening_nested_params + Enable flattening optional params to be used for custom layers. Use + this option if you want to transfer attributes of a custom layer to IR + with flattened nested parameters. Default behavior is to transfer + the attributes without flattening nested parameters. + + Kaldi-specific parameters: + --counts + Path to the counts file + --remove_output_softmax + Removes the SoftMax layer that is the output layer + --remove_memory + Removes the Memory layer and use additional inputs outputs instead + + + + +Fetching example models +----------------------- + +This notebook uses two models for conversion examples: + +- `Distilbert `__ + NLP model from Hugging Face +- `Resnet50 `__ + CV classification model from torchvision + +.. code:: ipython3 + + from pathlib import Path + + # create a directory for models files + MODEL_DIRECTORY_PATH = Path("model") + MODEL_DIRECTORY_PATH.mkdir(exist_ok=True) + +Fetch +`distilbert `__ +NLP model from Hugging Face and export it in ONNX format: + +.. code:: ipython3 + + from transformers import AutoModelForSequenceClassification, AutoTokenizer + from transformers.onnx import export, FeaturesManager + + + ONNX_NLP_MODEL_PATH = MODEL_DIRECTORY_PATH / "distilbert.onnx" + + # download model + hf_model = AutoModelForSequenceClassification.from_pretrained("distilbert-base-uncased-finetuned-sst-2-english") + # initialize tokenizer + tokenizer = AutoTokenizer.from_pretrained("distilbert-base-uncased-finetuned-sst-2-english") + + # get model onnx config function for output feature format sequence-classification + model_kind, model_onnx_config = FeaturesManager.check_supported_model_or_raise(hf_model, feature="sequence-classification") + # fill onnx config based on pytorch model config + onnx_config = model_onnx_config(hf_model.config) + + # export to onnx format + export(preprocessor=tokenizer, model=hf_model, config=onnx_config, opset=onnx_config.default_onnx_opset, output=ONNX_NLP_MODEL_PATH) + + +.. parsed-literal:: + + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/distilbert/modeling_distilbert.py:223: TracerWarning: torch.tensor results are registered as constants in the trace. You can safely ignore this warning if you use this function to create tensors out of constant variables that would be the same every time you call this function. In any other case, this might cause the trace to be incorrect. + mask, torch.tensor(torch.finfo(scores.dtype).min) + + + + +.. parsed-literal:: + + (['input_ids', 'attention_mask'], ['logits']) + + + +Fetch +`Resnet50 `__ +CV classification model from torchvision: + +.. code:: ipython3 + + from torchvision.models import resnet50, ResNet50_Weights + + + # create model object + pytorch_model = resnet50(weights=ResNet50_Weights.DEFAULT) + # switch model from training to inference mode + pytorch_model.eval() + + + + +.. parsed-literal:: + + ResNet( + (conv1): Conv2d(3, 64, kernel_size=(7, 7), stride=(2, 2), padding=(3, 3), bias=False) + (bn1): BatchNorm2d(64, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (relu): ReLU(inplace=True) + (maxpool): MaxPool2d(kernel_size=3, stride=2, padding=1, dilation=1, ceil_mode=False) + (layer1): Sequential( + (0): Bottleneck( + (conv1): Conv2d(64, 64, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn1): BatchNorm2d(64, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv2): Conv2d(64, 64, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), bias=False) + (bn2): BatchNorm2d(64, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv3): Conv2d(64, 256, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn3): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (relu): ReLU(inplace=True) + (downsample): Sequential( + (0): Conv2d(64, 256, kernel_size=(1, 1), stride=(1, 1), bias=False) + (1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + ) + ) + (1): Bottleneck( + (conv1): Conv2d(256, 64, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn1): BatchNorm2d(64, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv2): Conv2d(64, 64, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), bias=False) + (bn2): BatchNorm2d(64, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv3): Conv2d(64, 256, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn3): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (relu): ReLU(inplace=True) + ) + (2): Bottleneck( + (conv1): Conv2d(256, 64, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn1): BatchNorm2d(64, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv2): Conv2d(64, 64, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), bias=False) + (bn2): BatchNorm2d(64, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv3): Conv2d(64, 256, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn3): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (relu): ReLU(inplace=True) + ) + ) + (layer2): Sequential( + (0): Bottleneck( + (conv1): Conv2d(256, 128, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn1): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv2): Conv2d(128, 128, kernel_size=(3, 3), stride=(2, 2), padding=(1, 1), bias=False) + (bn2): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv3): Conv2d(128, 512, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn3): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (relu): ReLU(inplace=True) + (downsample): Sequential( + (0): Conv2d(256, 512, kernel_size=(1, 1), stride=(2, 2), bias=False) + (1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + ) + ) + (1): Bottleneck( + (conv1): Conv2d(512, 128, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn1): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv2): Conv2d(128, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), bias=False) + (bn2): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv3): Conv2d(128, 512, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn3): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (relu): ReLU(inplace=True) + ) + (2): Bottleneck( + (conv1): Conv2d(512, 128, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn1): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv2): Conv2d(128, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), bias=False) + (bn2): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv3): Conv2d(128, 512, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn3): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (relu): ReLU(inplace=True) + ) + (3): Bottleneck( + (conv1): Conv2d(512, 128, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn1): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv2): Conv2d(128, 128, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), bias=False) + (bn2): BatchNorm2d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv3): Conv2d(128, 512, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn3): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (relu): ReLU(inplace=True) + ) + ) + (layer3): Sequential( + (0): Bottleneck( + (conv1): Conv2d(512, 256, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv2): Conv2d(256, 256, kernel_size=(3, 3), stride=(2, 2), padding=(1, 1), bias=False) + (bn2): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv3): Conv2d(256, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (relu): ReLU(inplace=True) + (downsample): Sequential( + (0): Conv2d(512, 1024, kernel_size=(1, 1), stride=(2, 2), bias=False) + (1): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + ) + ) + (1): Bottleneck( + (conv1): Conv2d(1024, 256, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv2): Conv2d(256, 256, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), bias=False) + (bn2): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv3): Conv2d(256, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (relu): ReLU(inplace=True) + ) + (2): Bottleneck( + (conv1): Conv2d(1024, 256, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv2): Conv2d(256, 256, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), bias=False) + (bn2): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv3): Conv2d(256, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (relu): ReLU(inplace=True) + ) + (3): Bottleneck( + (conv1): Conv2d(1024, 256, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv2): Conv2d(256, 256, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), bias=False) + (bn2): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv3): Conv2d(256, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (relu): ReLU(inplace=True) + ) + (4): Bottleneck( + (conv1): Conv2d(1024, 256, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv2): Conv2d(256, 256, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), bias=False) + (bn2): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv3): Conv2d(256, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (relu): ReLU(inplace=True) + ) + (5): Bottleneck( + (conv1): Conv2d(1024, 256, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn1): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv2): Conv2d(256, 256, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), bias=False) + (bn2): BatchNorm2d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv3): Conv2d(256, 1024, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn3): BatchNorm2d(1024, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (relu): ReLU(inplace=True) + ) + ) + (layer4): Sequential( + (0): Bottleneck( + (conv1): Conv2d(1024, 512, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv2): Conv2d(512, 512, kernel_size=(3, 3), stride=(2, 2), padding=(1, 1), bias=False) + (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv3): Conv2d(512, 2048, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn3): BatchNorm2d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (relu): ReLU(inplace=True) + (downsample): Sequential( + (0): Conv2d(1024, 2048, kernel_size=(1, 1), stride=(2, 2), bias=False) + (1): BatchNorm2d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + ) + ) + (1): Bottleneck( + (conv1): Conv2d(2048, 512, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv2): Conv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), bias=False) + (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv3): Conv2d(512, 2048, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn3): BatchNorm2d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (relu): ReLU(inplace=True) + ) + (2): Bottleneck( + (conv1): Conv2d(2048, 512, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn1): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv2): Conv2d(512, 512, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), bias=False) + (bn2): BatchNorm2d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (conv3): Conv2d(512, 2048, kernel_size=(1, 1), stride=(1, 1), bias=False) + (bn3): BatchNorm2d(2048, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True) + (relu): ReLU(inplace=True) + ) + ) + (avgpool): AdaptiveAvgPool2d(output_size=(1, 1)) + (fc): Linear(in_features=2048, out_features=1000, bias=True) + ) + + + +Convert PyTorch model to ONNX format: + +.. code:: ipython3 + + import torch + import warnings + + + ONNX_CV_MODEL_PATH = MODEL_DIRECTORY_PATH / "resnet.onnx" + + if ONNX_CV_MODEL_PATH.exists(): + print(f"ONNX model {ONNX_CV_MODEL_PATH} already exists.") + else: + with warnings.catch_warnings(): + warnings.filterwarnings("ignore") + torch.onnx.export( + model=pytorch_model, + args=torch.randn(1, 3, 780, 520), + f=ONNX_CV_MODEL_PATH + ) + print(f"ONNX model exported to {ONNX_CV_MODEL_PATH}") + + +.. parsed-literal:: + + ONNX model exported to model/resnet.onnx + + +Basic conversion +---------------- + +To convert a model to OpenVINO IR, use the following command: + +.. code:: ipython3 + + # Model Optimizer CLI + + ! mo --input_model model/distilbert.onnx --output_dir model + + +.. parsed-literal:: + + huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks... + To disable this warning, you can either: + - Avoid using `tokenizers` before the fork if possible + - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false) + [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. + Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html + [ SUCCESS ] Generated IR version 11 model. + [ SUCCESS ] XML file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/distilbert.xml + [ SUCCESS ] BIN file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/distilbert.bin + + +.. code:: ipython3 + + # Python conversion API + from openvino.tools import mo + + # mo.convert_model returns an openvino.runtime.Model object + ov_model = mo.convert_model(ONNX_NLP_MODEL_PATH) + + # then model can be serialized to *.xml & *.bin files + from openvino.runtime import serialize + + serialize(ov_model, xml_path=MODEL_DIRECTORY_PATH / 'distilbert.xml') + + +.. parsed-literal:: + + huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks... + To disable this warning, you can either: + - Avoid using `tokenizers` before the fork if possible + - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false) + + +Model conversion parameters +--------------------------- + +Both Python conversion API and Model Optimizer command-line tool provide +the following capabilities: \* overriding original input shapes for +model conversion with ``input`` and ``input_shape`` parameters. `Setting +Input Shapes +guide `__. +\* cutting off unwanted parts of a model (such as unsupported operations +and training sub-graphs) using the ``input`` and ``output`` parameters +to define new inputs and outputs of the converted model. `Cutting Off +Parts of a Model +guide `__. +\* inserting additional input pre-processing sub-graphs into the +converted model by using the ``mean_values``, ``scales_values``, +``layout``, and other parameters. `Embedding Preprocessing Computation +article `__. +\* compressing the model weights (for example, weights for convolutions +and matrix multiplications) to FP16 data type using ``compress_to_fp16`` +compression parameter. `Compression of a Model to FP16 +guide `__. + +If the out-of-the-box conversion (only the ``input_model`` parameter is +specified) is not successful, it may be required to use the parameters +mentioned above to override input shapes and cut the model. + +Setting Input Shapes +~~~~~~~~~~~~~~~~~~~~ + +Model conversion is supported for models with dynamic input shapes that +contain undefined dimensions. However, if the shape of data is not going +to change from one inference request to another, it is recommended to +set up static shapes (when all dimensions are fully defined) for the +inputs. Doing it at this stage, instead of during inference in runtime, +can be beneficial in terms of performance and memory consumption. To set +up static shapes, model conversion API provides the ``input`` and +``input_shape`` parameters. + +For more information refer to `Setting Input Shapes +guide `__. + +.. code:: ipython3 + + # Model Optimizer CLI + + ! mo --input_model model/distilbert.onnx --input input_ids,attention_mask --input_shape [1,128],[1,128] --output_dir model + + # alternatively + ! mo --input_model model/distilbert.onnx --input input_ids[1,128],attention_mask[1,128] --output_dir model + + +.. parsed-literal:: + + huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks... + To disable this warning, you can either: + - Avoid using `tokenizers` before the fork if possible + - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false) + [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. + Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html + [ SUCCESS ] Generated IR version 11 model. + [ SUCCESS ] XML file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/distilbert.xml + [ SUCCESS ] BIN file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/distilbert.bin + huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks... + To disable this warning, you can either: + - Avoid using `tokenizers` before the fork if possible + - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false) + [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. + Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html + [ SUCCESS ] Generated IR version 11 model. + [ SUCCESS ] XML file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/distilbert.xml + [ SUCCESS ] BIN file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/distilbert.bin + + +.. code:: ipython3 + + # Python conversion API + from openvino.tools import mo + + + ov_model = mo.convert_model(ONNX_NLP_MODEL_PATH, input=["input_ids", "attention_mask"], input_shape=[[1, 128],[1, 128]]) + + # alternatively specify input shapes, using the input parameter + ov_model = mo.convert_model(ONNX_NLP_MODEL_PATH, input=[("input_ids", [1, 128]), ("attention_mask", [1, 128])]) + +The input_shape parameter allows overriding original input shapes to +ones compatible with a given model. Dynamic shapes, i.e. with dynamic +dimensions, can be replaced in the original model with static shapes for +the converted model, and vice versa. The dynamic dimension can be marked +in the model conversion API parameter as ``-1`` or ``?``. For example, +launch model conversion for the ONNX Bert model and specify a dynamic +sequence length dimension for inputs: + +.. code:: ipython3 + + # Model Optimizer CLI + + ! mo --input_model model/distilbert.onnx --input input_ids,attention_mask --input_shape [1,-1],[1,-1] --output_dir model + + +.. parsed-literal:: + + huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks... + To disable this warning, you can either: + - Avoid using `tokenizers` before the fork if possible + - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false) + [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. + Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html + [ SUCCESS ] Generated IR version 11 model. + [ SUCCESS ] XML file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/distilbert.xml + [ SUCCESS ] BIN file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/distilbert.bin + + +.. code:: ipython3 + + # Python conversion API + from openvino.tools import mo + + + ov_model = mo.convert_model(ONNX_NLP_MODEL_PATH, input=["input_ids", "attention_mask"], input_shape=[[1, -1],[1, -1]]) + +To optimize memory consumption for models with undefined dimensions in +runtime, model conversion API provides the capability to define +boundaries of dimensions. The boundaries of undefined dimensions can be +specified with ellipsis. For example, launch model conversion for the +ONNX Bert model and specify a boundary for the sequence length +dimension: + +.. code:: ipython3 + + # Model Optimizer CLI + + ! mo --input_model model/distilbert.onnx --input input_ids,attention_mask --input_shape [1,10..128],[1,10..128] --output_dir model + + +.. parsed-literal:: + + huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks... + To disable this warning, you can either: + - Avoid using `tokenizers` before the fork if possible + - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false) + [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. + Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html + [ SUCCESS ] Generated IR version 11 model. + [ SUCCESS ] XML file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/distilbert.xml + [ SUCCESS ] BIN file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/distilbert.bin + + +.. code:: ipython3 + + # Python conversion API + from openvino.tools import mo + + + ov_model = mo.convert_model(ONNX_NLP_MODEL_PATH, input=["input_ids", "attention_mask"], input_shape=[[1, "10..128"],[1, "10..128"]]) + +Cutting Off Parts of a Model +~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +The following examples show when model cutting is useful or even +required: + +- A model has pre- or post-processing parts that cannot be translated + to existing OpenVINO operations. +- A model has a training part that is convenient to be kept in the + model but not used during inference. +- A model is too complex to be converted at once because it contains + many unsupported operations that cannot be easily implemented as + custom layers. +- A problem occurs with model conversion or inference in OpenVINO + Runtime. To identify the issue, limit the conversion scope by an + iterative search for problematic areas in the model. +- A single custom layer or a combination of custom layers is isolated + for debugging purposes. + +For a more detailed description, refer to the `Cutting Off Parts of a +Model +guide `__. + +.. code:: ipython3 + + # Model Optimizer CLI + + # cut at the end + ! mo --input_model model/distilbert.onnx --output /classifier/Gemm --output_dir model + + + # cut from the beginning + ! mo --input_model model/distilbert.onnx --input /distilbert/embeddings/LayerNorm/Add_1,attention_mask --output_dir model + + +.. parsed-literal:: + + huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks... + To disable this warning, you can either: + - Avoid using `tokenizers` before the fork if possible + - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false) + [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. + Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html + [ SUCCESS ] Generated IR version 11 model. + [ SUCCESS ] XML file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/distilbert.xml + [ SUCCESS ] BIN file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/distilbert.bin + huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks... + To disable this warning, you can either: + - Avoid using `tokenizers` before the fork if possible + - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false) + [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. + Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html + [ SUCCESS ] Generated IR version 11 model. + [ SUCCESS ] XML file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/distilbert.xml + [ SUCCESS ] BIN file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/distilbert.bin + + +.. code:: ipython3 + + # Python conversion API + from openvino.tools import mo + + + # cut at the end + ov_model = mo.convert_model(ONNX_NLP_MODEL_PATH, output="/classifier/Gemm") + + # cut from the beginning + ov_model = mo.convert_model(ONNX_NLP_MODEL_PATH, input=["/distilbert/embeddings/LayerNorm/Add_1", "attention_mask"]) + +Embedding Preprocessing Computation +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +Input data for inference can be different from the training dataset and +requires additional preprocessing before inference. To accelerate the +whole pipeline, including preprocessing and inference, model conversion +API provides special parameters such as ``mean_values``, +``scale_values``, ``reverse_input_channels``, and ``layout``. Based on +these parameters, model conversion API generates OpenVINO IR with +additionally inserted sub-graphs to perform the defined preprocessing. +This preprocessing block can perform mean-scale normalization of input +data, reverting data along channel dimension, and changing the data +layout. For more information on preprocessing, refer to the `Embedding +Preprocessing Computation +article `__. + +Specifying Layout +^^^^^^^^^^^^^^^^^ + +Layout defines the meaning of dimensions in a shape and can be specified +for both inputs and outputs. Some preprocessing requires to set input +layouts, for example, setting a batch, applying mean or scales, and +reversing input channels (BGR<->RGB). For the layout syntax, check the +`Layout API +overview `__. +To specify the layout, you can use the layout option followed by the +layout value. + +The following command specifies the ``NCHW`` layout for a Pytorch +Resnet50 model that was exported to the ONNX format: + +.. code:: ipython3 + + # Model Optimizer CLI + + ! mo --input_model model/resnet.onnx --layout nchw --output_dir model + + +.. parsed-literal:: + + huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks... + To disable this warning, you can either: + - Avoid using `tokenizers` before the fork if possible + - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false) + [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. + Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html + [ SUCCESS ] Generated IR version 11 model. + [ SUCCESS ] XML file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/resnet.xml + [ SUCCESS ] BIN file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/resnet.bin + + +.. code:: ipython3 + + # Python conversion API + from openvino.tools import mo + + + ov_model = mo.convert_model(ONNX_CV_MODEL_PATH, layout="nchw") + +Changing Model Layout +^^^^^^^^^^^^^^^^^^^^^ + +Changing the model layout may be necessary if it differs from the one +presented by input data. Use either ``layout`` or ``source_layout`` with +``target_layout`` to change the layout. + +.. code:: ipython3 + + # Model Optimizer CLI + + ! mo --input_model model/resnet.onnx --layout "nchw->nhwc" --output_dir model + + # alternatively use source_layout and target_layout parameters + ! mo --input_model model/resnet.onnx --source_layout nchw --target_layout nhwc --output_dir model + + +.. parsed-literal:: + + huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks... + To disable this warning, you can either: + - Avoid using `tokenizers` before the fork if possible + - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false) + [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. + Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html + [ SUCCESS ] Generated IR version 11 model. + [ SUCCESS ] XML file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/resnet.xml + [ SUCCESS ] BIN file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/resnet.bin + huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks... + To disable this warning, you can either: + - Avoid using `tokenizers` before the fork if possible + - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false) + [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. + Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html + [ SUCCESS ] Generated IR version 11 model. + [ SUCCESS ] XML file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/resnet.xml + [ SUCCESS ] BIN file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/resnet.bin + + +.. code:: ipython3 + + # Python conversion API + from openvino.tools import mo + + + ov_model = mo.convert_model(ONNX_CV_MODEL_PATH, layout="nchw->nhwc") + + # alternatively use source_layout and target_layout parameters + ov_model = mo.convert_model(ONNX_CV_MODEL_PATH, source_layout="nchw", target_layout="nhwc") + +Specifying Mean and Scale Values +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +Model conversion API has the following parameters to specify the values: +``mean_values``, ``scale_values``, ``scale``. Using these parameters, +model conversion API embeds the corresponding preprocessing block for +mean-value normalization of the input data and optimizes this block so +that the preprocessing takes negligible time for inference. + +.. code:: ipython3 + + # Model Optimizer CLI + + ! mo --input_model model/resnet.onnx --mean_values [123,117,104] --scale 255 --output_dir model + + ! mo --input_model model/resnet.onnx --mean_values [123,117,104] --scale_values [255,255,255] --output_dir model + + +.. parsed-literal:: + + huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks... + To disable this warning, you can either: + - Avoid using `tokenizers` before the fork if possible + - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false) + [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. + Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html + [ SUCCESS ] Generated IR version 11 model. + [ SUCCESS ] XML file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/resnet.xml + [ SUCCESS ] BIN file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/resnet.bin + huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks... + To disable this warning, you can either: + - Avoid using `tokenizers` before the fork if possible + - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false) + [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. + Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html + [ SUCCESS ] Generated IR version 11 model. + [ SUCCESS ] XML file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/resnet.xml + [ SUCCESS ] BIN file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/resnet.bin + + +.. code:: ipython3 + + # Python conversion API + from openvino.tools import mo + + + ov_model = mo.convert_model(ONNX_CV_MODEL_PATH, mean_values=[123,117,104], scale=255) + + ov_model = mo.convert_model(ONNX_CV_MODEL_PATH, mean_values=[123,117,104], scale_values=[255,255,255]) + +Reversing Input Channels +^^^^^^^^^^^^^^^^^^^^^^^^ + +Sometimes, input images for your application can be of the ``RGB`` (or +``BGR``) format, and the model is trained on images of the ``BGR`` (or +``RGB``) format, which is in the opposite order of color channels. In +this case, it is important to preprocess the input images by reverting +the color channels before inference. + +.. code:: ipython3 + + # Model Optimizer CLI + + ! mo --input_model model/resnet.onnx --reverse_input_channels --output_dir model + + +.. parsed-literal:: + + huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks... + To disable this warning, you can either: + - Avoid using `tokenizers` before the fork if possible + - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false) + [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. + Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html + [ SUCCESS ] Generated IR version 11 model. + [ SUCCESS ] XML file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/resnet.xml + [ SUCCESS ] BIN file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/resnet.bin + + +.. code:: ipython3 + + # Python conversion API + from openvino.tools import mo + + + ov_model = mo.convert_model(ONNX_CV_MODEL_PATH, reverse_input_channels=True) + +Compressing a Model to FP16 +~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +Optionally all relevant floating-point weights can be compressed to FP16 +data type during the model conversion, creating a compressed FP16 model. +This smaller model occupies about half of the original space in the file +system. While the compression may introduce a drop in accuracy, for most +models, this decrease is negligible. + +.. code:: ipython3 + + # Model Optimizer CLI + + ! mo --input_model model/resnet.onnx --compress_to_fp16=True --output_dir model + + +.. parsed-literal:: + + huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks... + To disable this warning, you can either: + - Avoid using `tokenizers` before the fork if possible + - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false) + [ INFO ] Generated IR will be compressed to FP16. If you get lower accuracy, please consider disabling compression by removing argument --compress_to_fp16 or set it to false --compress_to_fp16=False. + Find more information about compression to FP16 at https://docs.openvino.ai/2023.0/openvino_docs_MO_DG_FP16_Compression.html + [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. + Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html + [ SUCCESS ] Generated IR version 11 model. + [ SUCCESS ] XML file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/resnet.xml + [ SUCCESS ] BIN file: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/121-convert-to-openvino/model/resnet.bin + + +.. code:: ipython3 + + # Python conversion API + from openvino.tools import mo + + + ov_model = mo.convert_model(ONNX_CV_MODEL_PATH, compress_to_fp16=True) + +Convert Models Represented as Python Objects +-------------------------------------------- + +Python conversion API can pass Python model objects, such as a Pytorch +model or TensorFlow Keras model directly, without saving them into files +and without leaving the training environment (Jupyter Notebook or +training scripts). + +.. code:: ipython3 + + # Python conversion API + from openvino.tools import mo + + + ov_model = mo.convert_model(pytorch_model) + +``convert_model()`` accepts all parameters available in the MO +command-line tool. Parameters can be specified by Python classes or +string analogs, similar to the command-line tool. + +.. code:: ipython3 + + # Python conversion API + from openvino.tools import mo + + + ov_model = mo.convert_model(pytorch_model, input_shape=[1,3,100,100], mean_values=[127, 127, 127], layout="nchw") + + ov_model = mo.convert_model(pytorch_model, source_layout="nchw", target_layout="nhwc") + + ov_model = mo.convert_model(pytorch_model, compress_to_fp16=True, reverse_input_channels=True) diff --git a/docs/notebooks/201-vision-monodepth-with-output.rst b/docs/notebooks/201-vision-monodepth-with-output.rst index d7ce6439d23..06ec0e5cd77 100644 --- a/docs/notebooks/201-vision-monodepth-with-output.rst +++ b/docs/notebooks/201-vision-monodepth-with-output.rst @@ -1,6 +1,8 @@ Monodepth Estimation with OpenVINO ================================== +.. _top: + This tutorial demonstrates Monocular Depth Estimation with MidasNet in OpenVINO. Model information can be found `here `__. @@ -26,13 +28,39 @@ Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer,” `__ in IEEE Transactions on Pattern Analysis and Machine Intelligence, doi: -10.1109/TPAMI.2020.3019967. +``10.1109/TPAMI.2020.3019967``. -Preparation ------------ +**Table of contents**: + +- `Preparation <#preparation>`__ + + - `Install requirements <#install-requirements>`__ + - `Imports <#imports>`__ + - `Download the model <#download-the-model>`__ + +- `Functions <#functions>`__ +- `Select inference device <#select-inference-device>`__ +- `Load the Model <#load-the-model>`__ +- `Monodepth on Image <#monodepth-on-image>`__ + + - `Load, resize and reshape input image <#load-resize-and-reshape-input-image>`__ + - `Do inference on the image <#do-inference-on-the-image>`__ + - `Display monodepth image <#display-monodepth-image>`__ + +- `Monodepth on Video <#monodepth-on-video>`__ + + - `Video Settings <#video-settings>`__ + - `Load the Video <#load-the-video>`__ + - `Do Inference on a Video and Create Monodepth Video <#do-inference-on-a-video-and-create-monodepth-video>`__ + - `Display Monodepth Video <#display-monodepth-video>`__ + +Preparation `⇑ <#top>`__ +############################################################################################################################### + + +Install requirements `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Install requirements -~~~~~~~~~~~~~~~~~~~~ .. code:: ipython3 @@ -51,12 +79,13 @@ Install requirements .. parsed-literal:: - ('notebook_utils.py', ) + ('notebook_utils.py', ) -Imports -~~~~~~~ +Imports `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -80,8 +109,9 @@ Imports from notebook_utils import download_file, load_image -Download the model -~~~~~~~~~~~~~~~~~~ +Download the model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -109,8 +139,9 @@ Download the model model/MiDaS_small.bin: 0%| | 0.00/31.6M [00:00`__ +############################################################################################################################### + .. code:: ipython3 @@ -142,10 +173,11 @@ Functions """ return cv2.cvtColor(image_data, cv2.COLOR_BGR2RGB) -Select inference device ------------------------ +Select inference device `⇑ <#top>`__ +############################################################################################################################### -select device from dropdown list for running inference using OpenVINO + +Select device from dropdown list for running inference using OpenVINO: .. code:: ipython3 @@ -170,8 +202,9 @@ select device from dropdown list for running inference using OpenVINO -Load the Model --------------- +Load the Model `⇑ <#top>`__ +############################################################################################################################### + Load the model in OpenVINO Runtime with ``ie.read_model`` and compile it for the specified device with ``ie.compile_model``. Get input and output @@ -190,11 +223,13 @@ keys and the expected input shape for the model. network_input_shape = list(input_key.shape) network_image_height, network_image_width = network_input_shape[2:] -Monodepth on Image ------------------- +Monodepth on Image `⇑ <#top>`__ +############################################################################################################################### + + +Load, resize and reshape input image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Load, resize and reshape input image -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The input image is read with OpenCV, resized to network input size, and reshaped to (N,C,H,W) (N=number of images, C=number of channels, @@ -211,8 +246,9 @@ H=height, W=width). # Reshape the image to network input shape NCHW. input_image = np.expand_dims(np.transpose(resized_image, (2, 0, 1)), 0) -Do inference on the image -~~~~~~~~~~~~~~~~~~~~~~~~~ +Do inference on the image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Do inference, convert the result to an image, and resize it to the original image shape. @@ -229,8 +265,9 @@ original image shape. # in (width, height), [::-1] reverses the (height, width) shape to match this. result_image = cv2.resize(result_image, image.shape[:2][::-1]) -Display monodepth image -~~~~~~~~~~~~~~~~~~~~~~~ +Display monodepth image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -243,15 +280,17 @@ Display monodepth image .. image:: 201-vision-monodepth-with-output_files/201-vision-monodepth-with-output_18_0.png -Monodepth on Video ------------------- +Monodepth on Video `⇑ <#top>`__ +############################################################################################################################### + By default, only the first 100 frames are processed in order to quickly check that everything works. Change ``NUM_FRAMES`` in the cell below to modify this. Set ``NUM_FRAMES`` to 0 to process the whole video. -Video Settings -~~~~~~~~~~~~~~ +Video Settings `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -279,8 +318,9 @@ Video Settings output_directory.mkdir(exist_ok=True) result_video_path = output_directory / f"{Path(VIDEO_FILE).stem}_monodepth.mp4" -Load the Video -~~~~~~~~~~~~~~ +Load the Video `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Load the video from a ``VIDEO_FILE``, set in the *Video Settings* cell above. Open the video to read the frame width and height and fps, and @@ -317,8 +357,9 @@ compute values for these properties for the monodepth video. The monodepth video will be scaled with a factor 0.5, have width 320, height 180, and run at 15.00 fps -Do Inference on a Video and Create Monodepth Video -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Do Inference on a Video and Create Monodepth Video `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -417,12 +458,13 @@ Do Inference on a Video and Create Monodepth Video .. parsed-literal:: - Processed 60 frames in 35.09 seconds. Total FPS (including video processing): 1.71.Inference FPS: 42.97 + Processed 60 frames in 48.08 seconds. Total FPS (including video processing): 1.25.Inference FPS: 42.27 Monodepth Video saved to 'output/Coco%20Walking%20in%20Berkeley_monodepth.mp4'. -Display Monodepth Video -~~~~~~~~~~~~~~~~~~~~~~~ +Display Monodepth Video `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -445,7 +487,7 @@ Display Monodepth Video .. parsed-literal:: Showing monodepth video saved at - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/201-vision-monodepth/output/Coco%20Walking%20in%20Berkeley_monodepth.mp4 + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/201-vision-monodepth/output/Coco%20Walking%20in%20Berkeley_monodepth.mp4 If you cannot see the video in your browser, please click on the following link to download the video diff --git a/docs/notebooks/201-vision-monodepth-with-output_files/index.html b/docs/notebooks/201-vision-monodepth-with-output_files/index.html index f72b319a48f..0f3de2b636b 100644 --- a/docs/notebooks/201-vision-monodepth-with-output_files/index.html +++ b/docs/notebooks/201-vision-monodepth-with-output_files/index.html @@ -1,7 +1,7 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/201-vision-monodepth-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/201-vision-monodepth-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/201-vision-monodepth-with-output_files/


../
-201-vision-monodepth-with-output_18_0.png          12-Jul-2023 00:11              959858
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/201-vision-monodepth-with-output_files/


../
+201-vision-monodepth-with-output_18_0.png          16-Aug-2023 01:31              959858
 

diff --git a/docs/notebooks/202-vision-superresolution-image-with-output.rst b/docs/notebooks/202-vision-superresolution-image-with-output.rst index c06e26dfd9e..c3902c024e8 100644 --- a/docs/notebooks/202-vision-superresolution-image-with-output.rst +++ b/docs/notebooks/202-vision-superresolution-image-with-output.rst @@ -1,6 +1,8 @@ Single Image Super Resolution with OpenVINO™ ============================================ +.. _top: + Super Resolution is the process of enhancing the quality of an image by increasing the pixel count using deep learning. This notebook shows the Single Image Super Resolution (SISR) which takes just one low resolution @@ -12,22 +14,56 @@ based on the research paper cited below. Y. Liu et al., `“An Attention-Based Approach for Single Image Super Resolution,” `__ 2018 24th International Conference on Pattern Recognition (ICPR), 2018, -pp. 2777-2784, doi: 10.1109/ICPR.2018.8545760. +pp. 2777-2784, doi: 10.1109/ICPR.2018.8545760. -Preparation ------------ +**Table of contents**: + +- `Preparation <#preparation>`__ + + - `Install requirements <#install-requirements>`__ + - `Imports <#imports>`__ + - `Settings <#settings>`__ + + - `Select inference device <#select-inference-device>`__ + + - `Functions <#functions>`__ + +- `Load the Superresolution Model <#load-the-superresolution-model>`__ +- `Load and Show the Input Image <#load-and-show-the-input-image>`__ +- `Superresolution on a Crop of the Image <#superresolution-on-a-crop-of-the-image>`__ + + - `Crop the Input Image once. <#crop-the-input-image-once>`__ + - `Reshape/Resize Crop for Model Input <#reshape-resize-crop-for-model-input>`__ + - `Do Inference <#do-inference>`__ + - `Show and Save Results <#show-and-save-results>`__ + + - `Save Superresolution and Bicubic Image Crop <#save-superresolution-and-bicubic-image-crop>`__ + - `Write Animated GIF with Bicubic/Superresolution Comparison <#write-animated-gif-with-bicubic-superresolution-comparison>`__ + - `Create a Video with Sliding Bicubic/Superresolution Comparison <#create-a-video-with-sliding-bicubic-superresolution-comparison>`__ + +- `Superresolution on full input image <#superresolution-on-full-input-image>`__ + + - `Compute patches <#compute-patches>`__ + - `Do Inference <#do-inference>`__ + - `Save superresolution image and the bicubic image <#save-superresolution-image-and-the-bicubic-image>`__ + +Preparation `⇑ <#top>`__ +############################################################################################################################### + + +Install requirements `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Install requirements -~~~~~~~~~~~~~~~~~~~~ .. code:: ipython3 - !pip install -q 'openvino>=2023.0.0' + !pip install -q "openvino>=2023.0.0" !pip install -q opencv-python !pip install -q pillow matplotlib -Imports -~~~~~~~ +Imports `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -53,13 +89,15 @@ Imports path.parent.mkdir(parents=True, exist_ok=True) urllib.request.urlretrieve(url, path) -Settings -~~~~~~~~ +Settings `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Select inference device -^^^^^^^^^^^^^^^^^^^^^^^ -select device from dropdown list for running inference using OpenVINO +Select inference device `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + + +Select device from dropdown list for running inference using OpenVINO: .. code:: ipython3 @@ -107,8 +145,9 @@ select device from dropdown list for running inference using OpenVINO else: print(f'{model_name} already downloaded to {base_model_dir}') -Functions -~~~~~~~~~ +Functions `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -168,8 +207,9 @@ Functions """ return cv2.cvtColor(image_data, cv2.COLOR_BGR2RGB) -Load the Superresolution Model ------------------------------- +Load the Superresolution Model `⇑ <#top>`__ +############################################################################################################################### + The Super Resolution model expects two inputs: the input image and a bicubic interpolation of the input image to the target size of @@ -216,12 +256,13 @@ about the network inputs and outputs. The image sides are upsampled by a factor of 4. The new image is 16 times as large as the original image -Load and Show the Input Image ------------------------------ +Load and Show the Input Image `⇑ <#top>`__ +############################################################################################################################### - **NOTE**: For the best results, use raw images (like TIFF, BMP or - PNG). Compressed images (like JPEG) may appear distorted after - processing with the super resolution model. + + **NOTE**: For the best results, use raw images (like ``TIFF``, + ``BMP`` or ``PNG``). Compressed images (like ``JPEG``) may appear + distorted after processing with the super resolution model. .. code:: ipython3 @@ -251,11 +292,13 @@ Load and Show the Input Image .. image:: 202-vision-superresolution-image-with-output_files/202-vision-superresolution-image-with-output_15_1.png -Superresolution on a Crop of the Image --------------------------------------- +Superresolution on a Crop of the Image `⇑ <#top>`__ +############################################################################################################################### + + +Crop the Input Image once. `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Crop the Input Image once. -~~~~~~~~~~~~~~~~~~~~~~~~~~ Crop the network input size. Give the X (width) and Y (height) coordinates for the top left corner of the crop. Set the ``CROP_FACTOR`` @@ -303,8 +346,9 @@ as the crop size. .. image:: 202-vision-superresolution-image-with-output_files/202-vision-superresolution-image-with-output_17_1.png -Reshape/Resize Crop for Model Input -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Reshape/Resize Crop for Model Input `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The input image is resized to a network input size, and reshaped to (N,C,H,W) (N=number of images, C=number of channels, H=height, W=width). @@ -326,8 +370,9 @@ interpolation. This bicubic image is the second input to the network. input_image_original = np.expand_dims(image_crop.transpose(2, 0, 1), axis=0) input_image_bicubic = np.expand_dims(bicubic_image.transpose(2, 0, 1), axis=0) -Do Inference -~~~~~~~~~~~~ +Do Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Do inference and convert the inference result to an ``RGB`` image. @@ -343,8 +388,9 @@ Do inference and convert the inference result to an ``RGB`` image. # Get inference result as numpy array and reshape to image shape and data type result_image = convert_result_to_image(result) -Show and Save Results -~~~~~~~~~~~~~~~~~~~~~ +Show and Save Results `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Show the bicubic image and the enhanced superresolution image. @@ -369,8 +415,9 @@ Show the bicubic image and the enhanced superresolution image. .. image:: 202-vision-superresolution-image-with-output_files/202-vision-superresolution-image-with-output_23_1.png -Save Superresolution and Bicubic Image Crop -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +Save Superresolution and Bicubic Image Crop `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + .. code:: ipython3 @@ -402,7 +449,7 @@ Save Superresolution and Bicubic Image Crop Write Animated GIF with Bicubic/Superresolution Comparison -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +`⇑ <#top>`__ .. code:: ipython3 @@ -440,7 +487,7 @@ Write Animated GIF with Bicubic/Superresolution Comparison Create a Video with Sliding Bicubic/Superresolution Comparison -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +`⇑ <#top>`__ This may take a while. For the video, the superresolution and bicubic image are resized by a factor of 2 to improve processing speed. This @@ -507,8 +554,9 @@ the ``Files`` tool. The video has been saved to output/flag_crop_comparison_2x.avi
-Superresolution on full input image ------------------------------------ +Superresolution on full input image `⇑ <#top>`__ +############################################################################################################################### + Superresolution on the full image is done by dividing the image into patches of equal size, doing superresolution on each path, and then @@ -518,8 +566,9 @@ near the border of the image are ignored. Adjust the ``CROPLINES`` setting in the next cell if you see boundary effects. -Compute patches -~~~~~~~~~~~~~~~ +Compute patches `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -563,8 +612,9 @@ Compute patches The output image will have a width of 11280 and a height of 7280 -Do Inference -~~~~~~~~~~~~ +Do Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The code below reads one patch of the image at a time. Each patch is reshaped to the network input shape and upsampled with bicubic @@ -681,12 +731,13 @@ as total time to process each patch. .. parsed-literal:: - Processed 42 patches in 4.73 seconds. Total patches per second (including processing): 8.87. - Inference patches per second: 17.65 + Processed 42 patches in 4.78 seconds. Total patches per second (including processing): 8.78. + Inference patches per second: 17.27 -Save superresolution image and the bicubic image -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Save superresolution image and the bicubic image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 diff --git a/docs/notebooks/202-vision-superresolution-image-with-output_files/index.html b/docs/notebooks/202-vision-superresolution-image-with-output_files/index.html index 6ac9a97ea11..321ed65740a 100644 --- a/docs/notebooks/202-vision-superresolution-image-with-output_files/index.html +++ b/docs/notebooks/202-vision-superresolution-image-with-output_files/index.html @@ -1,10 +1,10 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/202-vision-superresolution-image-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/202-vision-superresolution-image-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/202-vision-superresolution-image-with-output_files/


../
-202-vision-superresolution-image-with-output_15..> 12-Jul-2023 00:11              272963
-202-vision-superresolution-image-with-output_17..> 12-Jul-2023 00:11              356735
-202-vision-superresolution-image-with-output_23..> 12-Jul-2023 00:11             2896276
-202-vision-superresolution-image-with-output_27..> 12-Jul-2023 00:11             3207711
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/202-vision-superresolution-image-with-output_files/


../
+202-vision-superresolution-image-with-output_15..> 16-Aug-2023 01:31              272963
+202-vision-superresolution-image-with-output_17..> 16-Aug-2023 01:31              356735
+202-vision-superresolution-image-with-output_23..> 16-Aug-2023 01:31             2896276
+202-vision-superresolution-image-with-output_27..> 16-Aug-2023 01:31             3207711
 

diff --git a/docs/notebooks/202-vision-superresolution-video-with-output.rst b/docs/notebooks/202-vision-superresolution-video-with-output.rst index c7918bc1579..653f92d0d80 100644 --- a/docs/notebooks/202-vision-superresolution-video-with-output.rst +++ b/docs/notebooks/202-vision-superresolution-video-with-output.rst @@ -1,6 +1,8 @@ Video Super Resolution with OpenVINO™ ===================================== +.. _top: + Super Resolution is the process of enhancing the quality of an image by increasing the pixel count using deep learning. This notebook applies Single Image Super Resolution (SISR) to frames in a 360p (480×360) video @@ -16,24 +18,47 @@ pp. 2777-2784, doi: 10.1109/ICPR.2018.8545760. **NOTE**: The Single Image Super Resolution (SISR) model used in this demo is not optimized for a video. Results may vary depending on the - video. + video. -Preparation ------------ +**Table of contents**: -Imports -~~~~~~~ +- `Preparation <#preparation>`__ + + - `Install requirements <#install-requirements>`__ + - `Imports <#imports>`__ + - `Settings <#settings>`__ + + - `Select inference device <#select-inference-device>`__ + + - `Functions <#functions>`__ + +- `Load the Superresolution Model <#load-the-superresolution-model>`__ +- `Superresolution on Video <#superresolution-on-video>`__ + + - `Settings <#settings>`__ + - `Download and Prepare Video <#download-and-prepare-video>`__ + - `Do Inference <#do-inference>`__ + - `Show Side-by-Side Video of Bicubic and Superresolution Version <#show-side-by-side-video-of-bicubic-and-superresolution-version>`__ + +Preparation `⇑ <#top>`__ +############################################################################################################################### + +Install requirements `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ .. code:: ipython3 + !pip install -q "openvino>=2023.0.0" + !pip install -q opencv-python !pip install -q "pytube>=12.1.0" +Imports `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 import time - import requests from pathlib import Path - import sys import cv2 import numpy as np @@ -48,18 +73,49 @@ Imports ) from openvino.runtime import Core from pytube import YouTube - - sys.path.append("../utils") - from notebook_utils import download_file - -Settings -~~~~~~~~ .. code:: ipython3 - # Device to use for inference. For example, "CPU", or "GPU". - DEVICE = 'CPU' + # Define a download file helper function + def download_file(url: str, path: Path) -> None: + """Download file.""" + import urllib.request + path.parent.mkdir(parents=True, exist_ok=True) + urllib.request.urlretrieve(url, path) + +Settings `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + +Select inference device `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + core = Core() + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + # 1032: 4x superresolution, 1033: 3x superresolution model_name = 'single-image-super-resolution-1032' @@ -76,8 +132,8 @@ Settings model_xml_url = base_url + model_xml_name model_bin_url = base_url + model_bin_name - download_file(model_xml_url, model_xml_name, base_model_dir) - download_file(model_bin_url, model_bin_name, base_model_dir) + download_file(model_xml_url, model_xml_path) + download_file(model_bin_url, model_bin_path) else: print(f'{model_name} already downloaded to {base_model_dir}') @@ -87,68 +143,11 @@ Settings single-image-super-resolution-1032 already downloaded to model -Functions -~~~~~~~~~ +Functions `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ .. code:: ipython3 - def write_text_on_image(image: np.ndarray, text: str) -> np.ndarray: - """ - Write the specified text in the top left corner of the image - as white text with a black border. - - :param image: image as numpy array with HWC shape, RGB or BGR - :param text: text to write - :return: image with written text, as numpy array - """ - font = cv2.FONT_HERSHEY_PLAIN - org = (20, 20) - font_scale = 4 - font_color = (255, 255, 255) - line_type = 1 - font_thickness = 2 - text_color_bg = (0, 0, 0) - x, y = org - - image = cv2.UMat(image) - (text_w, text_h), _ = cv2.getTextSize( - text=text, fontFace=font, fontScale=font_scale, thickness=font_thickness - ) - result_im = cv2.rectangle( - img=image, pt1=org, pt2=(x + text_w, y + text_h), color=text_color_bg, thickness=-1 - ) - - textim = cv2.putText( - img=result_im, - text=text, - org=(x, y + text_h + font_scale - 1), - fontFace=font, - fontScale=font_scale, - color=font_color, - thickness=font_thickness, - lineType=line_type, - ) - return textim.get() - - - def load_image(path: str) -> np.ndarray: - """ - Loads an image from `path` and returns it as BGR numpy array. - - :param path: path to an image filename or url - :return: image as numpy array, with BGR channel order - """ - if path.startswith("http"): - # Set User-Agent to Mozilla because some websites block requests - # with User-Agent Python. - response = requests.get(url=path, headers={"User-Agent": "Mozilla/5.0"}) - array = np.asarray(bytearray(response.content), dtype="uint8") - image = cv2.imdecode(buf=array, flags=-1) # Loads the image as BGR. - else: - image = cv2.imread(filename=path) - return image - - def convert_result_to_image(result) -> np.ndarray: """ Convert network result of floating point numbers to image with integer @@ -163,17 +162,17 @@ Functions result = result.astype(np.uint8) return result -Load the Superresolution Model ------------------------------- +Load the Superresolution Model `⇑ <#top>`__ +############################################################################################################################### Load the model in OpenVINO Runtime with ``ie.read_model`` and compile it for the specified device with ``ie.compile_model``. .. code:: ipython3 - ie = Core() - model = ie.read_model(model=model_xml_path) - compiled_model = ie.compile_model(model=model, device_name=DEVICE) + core = Core() + model = core.read_model(model=model_xml_path) + compiled_model = core.compile_model(model=model, device_name=device.value) Get information about network inputs and outputs. The Super Resolution model expects two inputs: the input image and a bicubic interpolation of @@ -212,11 +211,11 @@ resolution version of the image in 1920x1080. The image sides are upsampled by a factor of 4. The new image is 16 times as large as the original image -Superresolution on Video ------------------------- +Superresolution on Video `⇑ <#top>`__ +############################################################################################################################### -Download a YouTube video with PyTube and enhance the video quality with -superresolution. +Download a YouTube video with ``PyTube`` and enhance the video quality +with superresolution. By default, only the first 100 frames of the video are processed. Change ``NUM_FRAMES`` in the cell below to modify this. @@ -225,12 +224,11 @@ By default, only the first 100 frames of the video are processed. Change should be a landscape video and have an input resolution of 360p (640x360) for the 1032 model, or 480p (720x480) for the 1033 model. -Settings -~~~~~~~~ +Settings `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ .. code:: ipython3 - VIDEO_DIR = "../data/video" OUTPUT_DIR = "output" Path(OUTPUT_DIR).mkdir(exist_ok=True) @@ -240,8 +238,8 @@ Settings # If you have FFMPEG installed, you can change FOURCC to `*"THEO"` to improve video writing speed. FOURCC = cv2.VideoWriter_fourcc(*"vp09") -Download and Prepare Video -~~~~~~~~~~~~~~~~~~~~~~~~~~ +Download and Prepare Video `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ .. code:: ipython3 @@ -252,18 +250,13 @@ Download and Prepare Video # Use `yt.streams` to see all available streams. See the PyTube documentation # https://python-pytube.readthedocs.io/en/latest/api.html for advanced # filtering options - try: - Path(VIDEO_DIR).mkdir(exist_ok=True) - stream = yt.streams.filter(resolution="360p").first() - filename = Path(stream.default_filename.encode("ascii", "ignore").decode("ascii")).stem - stream.download(output_path=OUTPUT_DIR, filename=filename) - print(f"Video {filename} downloaded to {OUTPUT_DIR}") + stream = yt.streams.filter(resolution="360p").first() + filename = Path(stream.default_filename.encode("ascii", "ignore").decode("ascii")).stem + stream.download(output_path=OUTPUT_DIR, filename=filename) + print(f"Video {filename} downloaded to {OUTPUT_DIR}") - # Create Path objects for the input video and the resulting videos. - video_path = Path(stream.get_file_path(filename, OUTPUT_DIR)) - except Exception: - # If PyTube fails, use a local video stored in the VIDEO_DIR directory. - video_path = Path(rf"{VIDEO_DIR}/CEO Pat Gelsinger on Leading Intel.mp4") + # Create Path objects for the input video and the resulting videos. + video_path = Path(stream.get_file_path(filename, OUTPUT_DIR)) # Path names for the result videos. superres_video_path = Path(f"{OUTPUT_DIR}/{video_path.stem}_superres.mp4") @@ -332,8 +325,8 @@ the superresolution side by side. frameSize=(target_width * 2, target_height), ) -Do Inference -~~~~~~~~~~~~ +Do Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ Read video frames and enhance them with superresolution. Save the superresolution video, the bicubic video and the comparison video to a @@ -449,17 +442,17 @@ video. .. parsed-literal:: - Processed frame 100. Inference time: 0.05 seconds (20.14 FPS) + Processed frame 100. Inference time: 0.06 seconds (16.26 FPS) .. parsed-literal:: Video's saved to output directory. - Processed 100 frames in 243.18 seconds. Total FPS (including video processing): 0.41. Inference FPS: 17.41. + Processed 100 frames in 235.27 seconds. Total FPS (including video processing): 0.43. Inference FPS: 17.68. -Show Side-by-Side Video of Bicubic and Superresolution Version -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Show Side-by-Side Video of Bicubic and Superresolution Version. `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ .. code:: ipython3 @@ -488,7 +481,7 @@ Show Side-by-Side Video of Bicubic and Superresolution Version .. raw:: html diff --git a/docs/notebooks/203-meter-reader-with-output.rst b/docs/notebooks/203-meter-reader-with-output.rst index b98677c3827..e45a6d9973c 100644 --- a/docs/notebooks/203-meter-reader-with-output.rst +++ b/docs/notebooks/203-meter-reader-with-output.rst @@ -1,6 +1,8 @@ Industrial Meter Reader ======================= +.. _top: + This notebook shows how to create a industrial meter reader with OpenVINO Runtime. We use the pre-trained `PPYOLOv2 `__ @@ -19,8 +21,26 @@ to build up a multiple inference task pipeline: workflow -Import ------- +**Table of contents**: + +- `Import <#import>`__ +- `Prepare the Model and Test Image <#prepare-the-model-and-test-image>`__ +- `Configuration <#configuration>`__ +- `Load the Models <#load-the-models>`__ +- `Data Process <#data-process>`__ +- `Main Function <#main-function>`__ + + - `Initialize the model and parameters. <#initialize-the-model-and-parameters>`__ + - `Run meter detection model <#run-meter-detection-model>`__ + - `Run meter segmentation model <#run-meter-segmentation-model>`__ + - `Postprocess the models result and calculate the final readings <#postprocess-the-models-result-and-calculate-the-final-readings>`__ + - `Get the reading result on the meter picture <#get-the-reading-result-on-the-meter-picture>`__ + +- `Try it with your meter photos! <#try-it-with-your-meter-photos>`__ + +Import `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -37,18 +57,17 @@ Import sys.path.append("../utils") from notebook_utils import download_file, segmentation_map_to_image -Prepare the Model and Test Image --------------------------------- +Prepare the Model and Test Image `⇑ <#top>`__ +############################################################################################################################### -Download PPYolov2 and DeepLabV3P pre-trained models from PaddlePaddle -community. +Download PPYOLOv2 and DeepLabV3P pre-trained models from PaddlePaddle community. .. code:: ipython3 MODEL_DIR = "model" DATA_DIR = "data" - DET_MODEL_LINK = "https://bj.bcebos.com/paddlex/examples2/meter_reader/meter_det_model.tar.gz" - SEG_MODEL_LINK = "https://bj.bcebos.com/paddlex/examples2/meter_reader/meter_seg_model.tar.gz" + DET_MODEL_LINK = "https://storage.openvinotoolkit.org/repositories/openvino_notebooks/models/meter-reader/meter_det_model.tar.gz" + SEG_MODEL_LINK = "https://storage.openvinotoolkit.org/repositories/openvino_notebooks/models/meter-reader/meter_seg_model.tar.gz" DET_FILE_NAME = DET_MODEL_LINK.split("/")[-1] SEG_FILE_NAME = SEG_MODEL_LINK.split("/")[-1] IMG_LINK = "https://user-images.githubusercontent.com/91237924/170696219-f68699c6-1e82-46bf-aaed-8e2fc3fa5f7b.jpg" @@ -113,8 +132,8 @@ community. Test Image Saved to "./data". -Configuration -------------- +Configuration `⇑ <#top>`__ +############################################################################################################################### Add parameter configuration for reading calculation. @@ -142,8 +161,8 @@ Add parameter configuration for reading calculation. SEG_LABEL = {'background': 0, 'pointer': 1, 'scale': 2} -Load the Models ---------------- +Load the Models `⇑ <#top>`__ +############################################################################################################################### Define a common class for model loading and inference @@ -185,8 +204,8 @@ Define a common class for model loading and inference result = self.compiled_model(input_image)[self.output_layer] return result -Data Process ------------- +Data Process `⇑ <#top>`__ +############################################################################################################################### Including the preprocessing and postprocessing tasks of each model. @@ -515,13 +534,15 @@ Including the preprocessing and postprocessing tasks of each model. readings.append(reading) return readings -Main Function -------------- +Main Function `⇑ <#top>`__ +############################################################################################################################### -Initialize the model and parameters. -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ -select device from dropdown list for running inference using OpenVINO +Initialize the model and parameters. `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: .. code:: ipython3 @@ -581,7 +602,7 @@ bounds of input batch size. .. parsed-literal:: - + @@ -589,11 +610,10 @@ bounds of input batch size. .. image:: 203-meter-reader-with-output_files/203-meter-reader-with-output_15_1.png -Run meter detection model -~~~~~~~~~~~~~~~~~~~~~~~~~ +Run meter detection model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Detect the location of the meter and prepare the ROI images for -segmentation. +Detect the location of the meter and prepare the ROI images for segmentation. .. code:: ipython3 @@ -634,8 +654,8 @@ segmentation. .. image:: 203-meter-reader-with-output_files/203-meter-reader-with-output_17_1.png -Run meter segmentation model -~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Run meter segmentation model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ Get the results of segmentation task on detected ROI. @@ -674,10 +694,11 @@ Get the results of segmentation task on detected ROI. .. image:: 203-meter-reader-with-output_files/203-meter-reader-with-output_19_1.png -Postprocess the models result and calculate the final readings -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Postprocess the models result and calculate the final readings `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Use OpenCV function to find the location of the pointer in a scale map. +Use OpenCV function to find the location of the pointer in a +scale map. .. code:: ipython3 @@ -711,8 +732,9 @@ Use OpenCV function to find the location of the pointer in a scale map. .. image:: 203-meter-reader-with-output_files/203-meter-reader-with-output_21_1.png -Get the reading result on the meter picture -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Get the reading result on the meter picture `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -742,4 +764,6 @@ Get the reading result on the meter picture .. image:: 203-meter-reader-with-output_files/203-meter-reader-with-output_23_1.png -## Try it with your meter photos! +Try it with your meter photos! `⇑ <#top>`__ +############################################################################################################################### + diff --git a/docs/notebooks/203-meter-reader-with-output_files/index.html b/docs/notebooks/203-meter-reader-with-output_files/index.html index 2dc0e89b32f..d2e984d5de6 100644 --- a/docs/notebooks/203-meter-reader-with-output_files/index.html +++ b/docs/notebooks/203-meter-reader-with-output_files/index.html @@ -1,11 +1,11 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/203-meter-reader-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/203-meter-reader-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/203-meter-reader-with-output_files/


../
-203-meter-reader-with-output_15_1.png              12-Jul-2023 00:11              170121
-203-meter-reader-with-output_17_1.png              12-Jul-2023 00:11              190271
-203-meter-reader-with-output_19_1.png              12-Jul-2023 00:11               26914
-203-meter-reader-with-output_21_1.png              12-Jul-2023 00:11                8966
-203-meter-reader-with-output_23_1.png              12-Jul-2023 00:11              170338
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/203-meter-reader-with-output_files/


../
+203-meter-reader-with-output_15_1.png              16-Aug-2023 01:31              170121
+203-meter-reader-with-output_17_1.png              16-Aug-2023 01:31              190271
+203-meter-reader-with-output_19_1.png              16-Aug-2023 01:31               26914
+203-meter-reader-with-output_21_1.png              16-Aug-2023 01:31                8966
+203-meter-reader-with-output_23_1.png              16-Aug-2023 01:31              170338
 

diff --git a/docs/notebooks/204-segmenter-semantic-segmentation-with-output.rst b/docs/notebooks/204-segmenter-semantic-segmentation-with-output.rst index 04b9833c30d..1db7b6e309c 100644 --- a/docs/notebooks/204-segmenter-semantic-segmentation-with-output.rst +++ b/docs/notebooks/204-segmenter-semantic-segmentation-with-output.rst @@ -1,6 +1,8 @@ Semantic Segmentation with OpenVINO™ using Segmenter ==================================================== +.. _top: + Semantic segmentation is a difficult computer vision problem with many applications such as autonomous driving, robotics, augmented reality, and many others. Its goal is to assign labels to each pixel according to @@ -24,7 +26,28 @@ Segmenter `__. More about the model and its details can be found in the following paper: `Segmenter: Transformer for Semantic Segmentation `__ or in the -`repository `__. +`repository `__. + +**Table of contents**: + +- `Get and prepare PyTorch model <#get-and-prepare-pytorch-model>`__ + + - `Prerequisites <#prerequisites>`__ + - `Loading PyTorch model <#loading-pytorch-model>`__ + +- `Preparing preprocessing and visualization functions <#preparing-preprocessing-and-visualization-functions>`__ + + - `Preprocessing <#preprocessing>`__ + - `Visualization <#visualization>`__ + +- `Validation of inference of original model <#validation-of-inference-of-original-model>`__ +- `Export to ONNX <#export-to-onnx>`__ +- `Convert ONNX model to OpenVINO Intermediate Representation (IR) <#convert-onnx-model-to-openvino-intermediate-representation-ir>`__ +- `Verify converted model inference <#verify-converted-model-inference>`__ + + - `Select inference device <#select-inference-device>`__ + +- `Benchmarking performance of converted model <#benchmarking-performance-of-converted-model>`__ .. |Segmenteer diagram| image:: https://user-images.githubusercontent.com/24582831/148507554-87eb80bd-02c7-4c31-b102-c6141e231ec8.png @@ -39,13 +62,14 @@ notebook consists of the following steps: - Validating inference of the converted model - Benchmark performance of the converted model -Get and prepare PyTorch model ------------------------------ +Get and prepare PyTorch model `⇑ <#top>`__ +############################################################################################################################### + The first thing we’ll need to do is clone `repository `__ containing model and helper functions. We will use Tiny model with mask transformer, that -is Seg-T-Mask/16. There are also better, but much larger models +is ``Seg-T-Mask/16``. There are also better, but much larger models available in the linked repo. This model is pre-trained on `ADE20K `__ dataset used for segmentation. @@ -54,8 +78,9 @@ The code from the repository already contains functions that create model and load weights, but we will need to download config and trained weights (checkpoint) file and add some additional helper functions. -Prerequisites -~~~~~~~~~~~~~ +Prerequisites `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -86,8 +111,8 @@ Prerequisites ) from notebook_utils import download_file, load_image -We’ll need timm, mmsegmentation, enops and mmcv, to use functions from -segmenter repo +We’ll need ``timm``, ``mmsegmentation``, ``einops`` and ``mmcv``, to use +functions from segmenter repo First, we will clone the Segmenter repo and then download weights and config for our model. @@ -109,7 +134,7 @@ config for our model. Cloning into 'segmenter'... remote: Enumerating objects: 268, done. remote: Total 268 (delta 0), reused 0 (delta 0), pack-reused 268 - Receiving objects: 100% (268/268), 15.34 MiB | 4.21 MiB/s, done. + Receiving objects: 100% (268/268), 15.34 MiB | 3.75 MiB/s, done. Resolving deltas: 100% (117/117), done. @@ -117,8 +142,8 @@ config for our model. # download config and pretrained model weights # here we use tiny model, there are also better but larger models available in repository - WEIGHTS_LINK = "https://www.rocq.inria.fr/cluster-willow/rstrudel/segmenter/checkpoints/ade20k/seg_tiny_mask/checkpoint.pth" - CONFIG_LINK = "https://www.rocq.inria.fr/cluster-willow/rstrudel/segmenter/checkpoints/ade20k/seg_tiny_mask/variant.yml" + WEIGHTS_LINK = "https://storage.openvinotoolkit.org/repositories/openvino_notebooks/models/segmenter/checkpoints/ade20k/seg_tiny_mask/checkpoint.pth" + CONFIG_LINK = "https://storage.openvinotoolkit.org/repositories/openvino_notebooks/models/segmenter/checkpoints/ade20k/seg_tiny_mask/variant.yml" MODEL_DIR = Path("model/") MODEL_DIR.mkdir(exist_ok=True) @@ -142,8 +167,9 @@ config for our model. model/variant.yml: 0%| | 0.00/940 [00:00`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + PyTorch models are usually an instance of `torch.nn.Module `__ @@ -187,26 +213,29 @@ Load normalization settings from config file. .. parsed-literal:: No CUDA runtime is found, using CUDA_HOME='/usr/local/cuda' - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/mmcv/__init__.py:20: UserWarning: On January 1, 2023, MMCV will release v2.0.0, in which it will remove components related to the training process and add a data transformation module. In addition, it will rename the package names mmcv to mmcv-lite and mmcv-full to mmcv. See https://github.com/open-mmlab/mmcv/blob/master/docs/en/compatibility.md for more details. + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/mmcv/__init__.py:20: UserWarning: On January 1, 2023, MMCV will release v2.0.0, in which it will remove components related to the training process and add a data transformation module. In addition, it will rename the package names mmcv to mmcv-lite and mmcv-full to mmcv. See https://github.com/open-mmlab/mmcv/blob/master/docs/en/compatibility.md for more details. warnings.warn( -Preparing preprocessing and visualization functions ---------------------------------------------------- +Preparing preprocessing and visualization functions `⇑ <#top>`__ +############################################################################################################################### + Now we will define utility functions for preprocessing and visualizing the results. -Preprocessing -~~~~~~~~~~~~~ +Preprocessing `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Inference input is tensor with shape ``[1, 3, H, W]`` in ``B, C, H, W`` format, where: -* ``B`` - batch size (in our case 1, as we are just adding 1 with unsqueeze) -* ``C`` - image channels (in our case RGB - 3) -* ``H`` - image height -* ``W`` - image width +- ``B`` - batch size (in our case 1, as we are just adding 1 with + unsqueeze) +- ``C`` - image channels (in our case RGB - 3) +- ``H`` - image height +- ``W`` - image width Resizing to the correct scale and splitting to batches is done inside inference, so we don’t need to resize or split the image in @@ -241,15 +270,16 @@ normalized with given mean and standard deviation provided in return im -Visualization -~~~~~~~~~~~~~ +Visualization `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Inference output contains labels assigned to each pixel, so the output -in our case is ``[150, H, W]`` in ``CL, H, W`` format where: +in our case is ``[150, H, W]`` in ``CL, H, W`` format where: -* ``CL`` - number of classes for labels (in our case 150) -* ``H`` - image height -* ``W`` - image width +- ``CL`` - number of classes for labels (in our case 150) +- ``H`` - image height +- ``W`` - image width Since we want to visualize this output, we reduce dimensions to ``[1, H, W]`` where we keep only class with the highest value as that is @@ -285,8 +315,9 @@ corresponding to the inferred labels. return pil_blend -Validation of inference of original model ------------------------------------------ +Validation of inference of original model `⇑ <#top>`__ +############################################################################################################################### + Now that we have everything ready, we can perform segmentation on example image ``coco_hollywood.jpg``. @@ -337,8 +368,9 @@ We can see that model segments the image into meaningful parts. Since we are using tiny variant of model, the result is not as good as it is with larger models, but it already shows nice segmentation performance. -Export to ONNX --------------- +Export to ONNX `⇑ <#top>`__ +############################################################################################################################### + Now that we’ve verified that the inference of PyTorch model works, we will first export it to ONNX format. @@ -347,15 +379,16 @@ To do this, we first get input dimensions from the model configuration file and create torch dummy input. Input dimensions are in our case ``[2, 3, 512, 512]`` in ``B, C, H, W]`` format, where: -* ``B`` - batch size -* ``C`` - image channels (in our case RGB - 3) -* ``H`` - model input image height -* ``W`` - model input image width +- ``B`` - batch size +- ``C`` - image channels (in our case RGB - 3) +- ``H`` - model input image height +- ``W`` - model input image width -.. note:: +.. - Note that H and W are here fixed to 512, as this is required by the model. Resizing is -done inside the inference function from the original repository. + Note that H and W are here fixed to 512, as this is required by the + model. Resizing is done inside the inference function from the + original repository. After that, we use ``export`` function from PyTorch to convert the model to ONNX. The process can generate some warnings, but they are not a @@ -388,36 +421,37 @@ problem. .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/204-segmenter-semantic-segmentation/./segmenter/segm/model/utils.py:69: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/204-segmenter-semantic-segmentation/./segmenter/segm/model/utils.py:69: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if H % patch_size > 0: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/204-segmenter-semantic-segmentation/./segmenter/segm/model/utils.py:71: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/204-segmenter-semantic-segmentation/./segmenter/segm/model/utils.py:71: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if W % patch_size > 0: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/204-segmenter-semantic-segmentation/./segmenter/segm/model/vit.py:122: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/204-segmenter-semantic-segmentation/./segmenter/segm/model/vit.py:122: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if x.shape[1] != pos_embed.shape[1]: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/204-segmenter-semantic-segmentation/./segmenter/segm/model/decoder.py:100: TracerWarning: Converting a tensor to a Python integer might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/204-segmenter-semantic-segmentation/./segmenter/segm/model/decoder.py:100: TracerWarning: Converting a tensor to a Python integer might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! masks = rearrange(masks, "b (h w) n -> b n h w", h=int(GS)) - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/204-segmenter-semantic-segmentation/./segmenter/segm/model/utils.py:85: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/204-segmenter-semantic-segmentation/./segmenter/segm/model/utils.py:85: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if extra_h > 0: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/204-segmenter-semantic-segmentation/./segmenter/segm/model/utils.py:87: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/204-segmenter-semantic-segmentation/./segmenter/segm/model/utils.py:87: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if extra_w > 0: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/_internal/jit_utils.py:258: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/_internal/jit_utils.py:258: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) _C._jit_pass_onnx_node_shape_type_inference(node, params_dict, opset_version) - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) _C._jit_pass_onnx_graph_shape_type_inference( - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) _C._jit_pass_onnx_graph_shape_type_inference( -Convert ONNX model to OpenVINO Intermediate Representation (IR) ---------------------------------------------------------------- +Convert ONNX model to OpenVINO Intermediate Representation (IR). `⇑ <#top>`__ +############################################################################################################################### While ONNX models are directly supported by OpenVINO runtime, it can be useful to convert them to IR format to take advantage of OpenVINO -optimization tools and features. ``mo.convert_model`` function can be -used for converting model using OpenVINO Model Optimizer. The function -returns instance of OpenVINO Model class, which is ready to use in -Python interface but can also be serialized to OpenVINO IR format for -future execution. +optimization tools and features. The ``mo.convert_model`` function of +`model conversion +API `__ +can be used. The function returns instance of OpenVINO Model class, +which is ready to use in Python interface but can also be serialized to +OpenVINO IR format for future execution. .. code:: ipython3 @@ -428,8 +462,9 @@ future execution. # serialize model for saving IR serialize(model, str(MODEL_DIR / "segmenter.xml")) -Verify converted model inference --------------------------------- +Verify converted model inference `⇑ <#top>`__ +############################################################################################################################### + To test that model was successfully converted, we can use same inference function from original repository, but we need to make custom class. @@ -496,13 +531,14 @@ any additional custom code required to process input. """ return torch.from_numpy(self.model(data)[self.output_blob]) -Now that we have created SegmenterOV helper class, we can use it in +Now that we have created ``SegmenterOV`` helper class, we can use it in inference function. -Select inference device -~~~~~~~~~~~~~~~~~~~~~~~ +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -select device from dropdown list for running inference using OpenVINO + +Select device from dropdown list for running inference using OpenVINO: .. code:: ipython3 @@ -560,8 +596,9 @@ select device from dropdown list for running inference using OpenVINO As we can see, we get the same results as with original model. -Benchmarking performance of converted model -------------------------------------------- +Benchmarking performance of converted model `⇑ <#top>`__ +############################################################################################################################### + Finally, use the OpenVINO `Benchmark Tool `__ @@ -606,18 +643,18 @@ to measure the inference performance of the model. [Step 2/11] Loading OpenVINO Runtime [ WARNING ] Default duration 120 seconds is used for unknown device AUTO [ INFO ] OpenVINO: - [ INFO ] Build ................................. 2023.0.0-10926-b4452d56304-releases/2023/0 + [ INFO ] Build ................................. 2023.0.1-11005-fa1c41994f3-releases/2023/0 [ INFO ] [ INFO ] Device info: [ INFO ] AUTO - [ INFO ] Build ................................. 2023.0.0-10926-b4452d56304-releases/2023/0 + [ INFO ] Build ................................. 2023.0.1-11005-fa1c41994f3-releases/2023/0 [ INFO ] [ INFO ] [Step 3/11] Setting device configuration [ WARNING ] Performance hint was not explicitly specified in command line. Device(AUTO) performance hint will be set to PerformanceMode.THROUGHPUT. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 16.16 ms + [ INFO ] Read model took 16.30 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] input (node: input) : f32 / [...] / [2,3,512,512] @@ -631,7 +668,7 @@ to measure the inference performance of the model. [ INFO ] Model outputs: [ INFO ] output (node: output) : f32 / [...] / [2,150,512,512] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 349.84 ms + [ INFO ] Compile model took 343.28 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT @@ -660,15 +697,15 @@ to measure the inference performance of the model. [ INFO ] Fill input 'input' with random values [Step 10/11] Measuring performance (Start inference asynchronously, 6 inference requests, limits: 120000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 219.27 ms + [ INFO ] First inference took 227.83 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 1374 iterations - [ INFO ] Duration: 120364.60 ms + [ INFO ] Count: 1332 iterations + [ INFO ] Duration: 120630.17 ms [ INFO ] Latency: - [ INFO ] Median: 523.66 ms - [ INFO ] Average: 524.91 ms - [ INFO ] Min: 217.86 ms - [ INFO ] Max: 612.22 ms - [ INFO ] Throughput: 22.83 FPS + [ INFO ] Median: 542.28 ms + [ INFO ] Average: 542.75 ms + [ INFO ] Min: 344.15 ms + [ INFO ] Max: 609.17 ms + [ INFO ] Throughput: 22.08 FPS diff --git a/docs/notebooks/204-segmenter-semantic-segmentation-with-output_files/index.html b/docs/notebooks/204-segmenter-semantic-segmentation-with-output_files/index.html index b085b8d768f..bb43538de1d 100644 --- a/docs/notebooks/204-segmenter-semantic-segmentation-with-output_files/index.html +++ b/docs/notebooks/204-segmenter-semantic-segmentation-with-output_files/index.html @@ -1,10 +1,10 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/204-segmenter-semantic-segmentation-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/204-segmenter-semantic-segmentation-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/204-segmenter-semantic-segmentation-with-output_files/


../
-204-segmenter-semantic-segmentation-with-output..> 12-Jul-2023 00:11               72352
-204-segmenter-semantic-segmentation-with-output..> 12-Jul-2023 00:11              909669
-204-segmenter-semantic-segmentation-with-output..> 12-Jul-2023 00:11               72356
-204-segmenter-semantic-segmentation-with-output..> 12-Jul-2023 00:11              909691
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/204-segmenter-semantic-segmentation-with-output_files/


../
+204-segmenter-semantic-segmentation-with-output..> 16-Aug-2023 01:31               72352
+204-segmenter-semantic-segmentation-with-output..> 16-Aug-2023 01:31              909669
+204-segmenter-semantic-segmentation-with-output..> 16-Aug-2023 01:31               72356
+204-segmenter-semantic-segmentation-with-output..> 16-Aug-2023 01:31              909691
 

diff --git a/docs/notebooks/205-vision-background-removal-with-output.rst b/docs/notebooks/205-vision-background-removal-with-output.rst index 51688a4e9cc..cd53815c483 100644 --- a/docs/notebooks/205-vision-background-removal-with-output.rst +++ b/docs/notebooks/205-vision-background-removal-with-output.rst @@ -1,24 +1,51 @@ Image Background Removal with U^2-Net and OpenVINO™ =================================================== +.. _top: + This notebook demonstrates background removal in images using U\ :math:`^2`-Net and OpenVINO. For more information about U\ :math:`^2`-Net, including source code and -test data, see the `Github +test data, see the `GitHub page `__ and the research paper: `U^2-Net: Going Deeper with Nested U-Structure for Salient Object Detection `__. The PyTorch U\ :math:`^2`-Net model is converted to OpenVINO IR format. The model source is available -`here `__. +`here `__. -Preparation ------------ -Install requirements -~~~~~~~~~~~~~~~~~~~~ +**Table of contents**: + +- `Preparation <#preparation>`__ + + - `Install requirements <#install-requirements>`__ + - `Import the PyTorch Library and U2-Net <#import-the-pytorch-library-and-u2-net>`__ + - `Settings <#settings>`__ + - `Load the U2-Net Model <#load-the-u2-net-model>`__ + +- `Convert PyTorch U2-Net model to OpenVINO IR <#convert-pytorch-u2-net-model-to-openvino-ir>`__ + + - `Convert Pytorch model to OpenVINO IR Format <#convert-pytorch-model-to-openvino-ir-format>`__ + +- `Load and Pre-Process Input Image <#load-and-pre-process-input-image>`__ +- `Select inference device <#select-inference-device>`__ +- `Do Inference on OpenVINO IR Model <#do-inference-on-openvino-ir-model>`__ +- `Visualize Results <#visualize-results>`__ + + - `Add a Background Image <#add-a-background-image>`__ + +- `References <#references>`__ + +Preparation `⇑ <#top>`__ +############################################################################################################################### + + +Install requirements `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -26,8 +53,9 @@ Install requirements !pip install -q torch onnx opencv-python matplotlib !pip install -q gdown -Import the PyTorch Library and U\ :math:`^2`-Net -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Import the PyTorch Library and U\ :math:`^2`-Net `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -63,8 +91,9 @@ Import the PyTorch Library and U\ :math:`^2`-Net from notebook_utils import load_image from model.u2net import U2NET, U2NETP -Settings -~~~~~~~~ +Settings `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + This tutorial supports using the original U\ :math:`^2`-Net salient object detection model, as well as the smaller U2NETP version. Two sets @@ -103,8 +132,9 @@ detection and human segmentation. MODEL_DIR = "model" model_path = Path(MODEL_DIR) / u2net_model.name / Path(u2net_model.name).with_suffix(".pth") -Load the U\ :math:`^2`-Net Model -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Load the U\ :math:`^2`-Net Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The U\ :math:`^2`-Net human segmentation model weights are stored on Google Drive. They will be downloaded if they are not present yet. The @@ -132,8 +162,7 @@ next cell loads the model and the pre-trained weights. Downloading... From: https://drive.google.com/uc?id=1rbSTGKAE-MTxBYHd-51l2hMOQPT_7EPy To: <_io.BufferedWriter name='model/u2net_lite/u2net_lite.pth'> - 100%|██████████| 4.68M/4.68M [00:00<00:00, 4.90MB/s] - + 100%|██████████| 4.68M/4.68M [00:01<00:00, 3.98MB/s] .. parsed-literal:: @@ -160,37 +189,39 @@ next cell loads the model and the pre-trained weights. .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/nn/functional.py:3734: UserWarning: nn.functional.upsample is deprecated. Use nn.functional.interpolate instead. + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/nn/functional.py:3734: UserWarning: nn.functional.upsample is deprecated. Use nn.functional.interpolate instead. warnings.warn("nn.functional.upsample is deprecated. Use nn.functional.interpolate instead.") - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/nn/functional.py:1967: UserWarning: nn.functional.sigmoid is deprecated. Use torch.sigmoid instead. + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/nn/functional.py:1967: UserWarning: nn.functional.sigmoid is deprecated. Use torch.sigmoid instead. warnings.warn("nn.functional.sigmoid is deprecated. Use torch.sigmoid instead.") - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/_internal/jit_utils.py:258: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/_internal/jit_utils.py:258: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) _C._jit_pass_onnx_node_shape_type_inference(node, params_dict, opset_version) - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) _C._jit_pass_onnx_graph_shape_type_inference( - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) _C._jit_pass_onnx_graph_shape_type_inference( -Convert PyTorch U\ :math:`^2`-Net model to OpenVINO IR ------------------------------------------------------- +Convert PyTorch U\ :math:`^2`-Net model to OpenVINO IR `⇑ <#top>`__ +############################################################################################################################### -Convert Pytorch model to OpenVINO IR Format -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ -Use Model Optimizer Python API to convert the Pytorch model to OpenVINO -IR format, with ``FP16`` precision. We add the mean values to the model -and scale the input with the standard deviation with ``scale_values`` -parameter. With these options, it is not necessary to normalize input -data before propagating it through the network. The mean and standard -deviation values can be found in the +Convert Pytorch model to OpenVINO IR Format `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +To convert the Pytorch model to OpenVINO IR format with ``FP16`` +precision, use model conversion Python API . We add the mean values to +the model and scale the input with the standard deviation with +``scale_values`` parameter. With these options, it is not necessary to +normalize input data before propagating it through the network. The mean +and standard deviation values can be found in the `dataloader `__ file in the `U^2-Net repository `__ and multiplied by 255 to support images with pixel values from 0-255. -For more information, refer to the `Model Optimizer Developer -Guide `__. +For more information about model conversion, refer to this +`page `__. Executing the following command may take a while. @@ -203,8 +234,9 @@ Executing the following command may take a while. compress_to_fp16=True ) -Load and Pre-Process Input Image --------------------------------- +Load and Pre-Process Input Image `⇑ <#top>`__ +############################################################################################################################### + While OpenCV reads images in ``BGR`` format, the OpenVINO IR model expects images in ``RGB``. Therefore, convert the images to ``RGB``, @@ -224,16 +256,46 @@ that is expected by the OpenVINO IR model. # for OpenVINO IR model: (1, 3, 512, 512). input_image = np.expand_dims(np.transpose(resized_image, (2, 0, 1)), 0) -Do Inference on OpenVINO IR Model ---------------------------------- +Select inference device `⇑ <#top>`__ +############################################################################################################################### + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + core = Core() + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +Do Inference on OpenVINO IR Model `⇑ <#top>`__ +############################################################################################################################### + Load the OpenVINO IR model to OpenVINO Runtime and do inference. .. code:: ipython3 # Load the network to OpenVINO Runtime. - ie = Core() - compiled_model_ir = ie.compile_model(model=model_ir, device_name="CPU") + core = Core() + compiled_model_ir = core.compile_model(model=model_ir, device_name=device.value) # Get the names of input and output layers. input_layer_ir = compiled_model_ir.input(0) output_layer_ir = compiled_model_ir.output(0) @@ -250,11 +312,12 @@ Load the OpenVINO IR model to OpenVINO Runtime and do inference. .. parsed-literal:: - Inference finished. Inference time: 0.122 seconds, FPS: 8.17. + Inference finished. Inference time: 0.119 seconds, FPS: 8.43. -Visualize Results ------------------ +Visualize Results `⇑ <#top>`__ +############################################################################################################################### + Show the original image, the segmentation result, and the original image with the background removed. @@ -281,11 +344,12 @@ with the background removed. -.. image:: 205-vision-background-removal-with-output_files/205-vision-background-removal-with-output_20_0.png +.. image:: 205-vision-background-removal-with-output_files/205-vision-background-removal-with-output_22_0.png -Add a Background Image -~~~~~~~~~~~~~~~~~~~~~~ +Add a Background Image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + In the segmentation result, all foreground pixels have a value of 1, all background pixels a value of 0. Replace the background image as follows: @@ -340,7 +404,7 @@ background pixels a value of 0. Replace the background image as follows: -.. image:: 205-vision-background-removal-with-output_files/205-vision-background-removal-with-output_22_0.png +.. image:: 205-vision-background-removal-with-output_files/205-vision-background-removal-with-output_24_0.png @@ -349,13 +413,14 @@ background pixels a value of 0. Replace the background image as follows: The generated image coco_hollywood-wall.jpg is saved in the directory output. You can also download the image by clicking on this link: output/coco_hollywood-wall.jpg
-References ----------- +References `⇑ <#top>`__ +############################################################################################################################### + - `PIP install openvino-dev `__ -- `Model Optimizer - Documentation `__ +- `Model Conversion + API `__ - `U^2-Net `__ - U^2-Net research paper: `U^2-Net: Going Deeper with Nested U-Structure for Salient Object diff --git a/docs/notebooks/205-vision-background-removal-with-output_files/205-vision-background-removal-with-output_20_0.png b/docs/notebooks/205-vision-background-removal-with-output_files/205-vision-background-removal-with-output_20_0.png deleted file mode 100644 index e56801a4118..00000000000 --- a/docs/notebooks/205-vision-background-removal-with-output_files/205-vision-background-removal-with-output_20_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:83192799a2c31d56beff242e7d9b1dae795e6cb3ef9c6f754c8cdc102f3af8dd -size 279567 diff --git a/docs/notebooks/205-vision-background-removal-with-output_files/205-vision-background-removal-with-output_22_0.png b/docs/notebooks/205-vision-background-removal-with-output_files/205-vision-background-removal-with-output_22_0.png index 8fd15b011dc..e56801a4118 100644 --- a/docs/notebooks/205-vision-background-removal-with-output_files/205-vision-background-removal-with-output_22_0.png +++ b/docs/notebooks/205-vision-background-removal-with-output_files/205-vision-background-removal-with-output_22_0.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:a63381a3ccf5d7a134519b437c3197abb82df80edc51bc8fa51c574a186ac1cc -size 927148 +oid sha256:83192799a2c31d56beff242e7d9b1dae795e6cb3ef9c6f754c8cdc102f3af8dd +size 279567 diff --git a/docs/notebooks/205-vision-background-removal-with-output_files/205-vision-background-removal-with-output_24_0.png b/docs/notebooks/205-vision-background-removal-with-output_files/205-vision-background-removal-with-output_24_0.png new file mode 100644 index 00000000000..8fd15b011dc --- /dev/null +++ b/docs/notebooks/205-vision-background-removal-with-output_files/205-vision-background-removal-with-output_24_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:a63381a3ccf5d7a134519b437c3197abb82df80edc51bc8fa51c574a186ac1cc +size 927148 diff --git a/docs/notebooks/205-vision-background-removal-with-output_files/index.html b/docs/notebooks/205-vision-background-removal-with-output_files/index.html index fa034c30579..c264929e7de 100644 --- a/docs/notebooks/205-vision-background-removal-with-output_files/index.html +++ b/docs/notebooks/205-vision-background-removal-with-output_files/index.html @@ -1,8 +1,8 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/205-vision-background-removal-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/205-vision-background-removal-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/205-vision-background-removal-with-output_files/


../
-205-vision-background-removal-with-output_20_0.png 12-Jul-2023 00:11              279567
-205-vision-background-removal-with-output_22_0.png 12-Jul-2023 00:11              927148
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/205-vision-background-removal-with-output_files/


../
+205-vision-background-removal-with-output_22_0.png 16-Aug-2023 01:31              279567
+205-vision-background-removal-with-output_24_0.png 16-Aug-2023 01:31              927148
 

diff --git a/docs/notebooks/206-vision-paddlegan-anime-with-output.rst b/docs/notebooks/206-vision-paddlegan-anime-with-output.rst index f3063b7623e..da7002cc9e9 100644 --- a/docs/notebooks/206-vision-paddlegan-anime-with-output.rst +++ b/docs/notebooks/206-vision-paddlegan-anime-with-output.rst @@ -1,6 +1,8 @@ Photos to Anime with PaddleGAN and OpenVINO =========================================== +.. _top: + This tutorial demonstrates converting a `PaddlePaddle/PaddleGAN `__ AnimeGAN model to OpenVINO IR format, and shows inference results on the @@ -14,17 +16,45 @@ documentation `__ + + - `Install requirements <#install-requirements>`__ + - `Imports <#imports>`__ + - `Settings <#settings>`__ + - `Functions <#functions>`__ + +- `Inference on PaddleGAN Model <#inference-on-paddlegan-model>`__ + + - `Show Inference Results on PaddleGAN model <#show-inference-results-on-paddlegan-model>`__ + +- `Model Conversion to ONNX and OpenVINO IR <#model-conversion-to-onnx-and-openvino-ir>`__ + + - `Convert to ONNX <#convert-to-onnx>`__ + - `Convert to OpenVINO IR <#convert-to-openvino-ir>`__ + +- `Show Inference Results on OpenVINO IR and PaddleGAN Models <#show-inference-results-on-openvino-ir-and-paddlegan-models>`__ + + - `Create Postprocessing Functions <#create-postprocessing-functions>`__ + - `Do Inference on OpenVINO IR Model <#do-inference-on-openvino-ir-model>`__ + + - `Select inference device <#select-inference-device>`__ + +- `Performance Comparison <#performance-comparison>`__ +- `References <#references>`__ + +Preparation `⇑ <#top>`__ +############################################################################################################################### + +Install requirements `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ .. code:: ipython3 !pip install -q "openvino-dev>=2023.0.0" - !pip install -q "paddlepaddle==2.5.0rc0" "paddle2onnx>=0.6" + !pip install -q "paddlepaddle==2.5.0" "paddle2onnx>=0.6" !pip install -q "git+https://github.com/PaddlePaddle/PaddleGAN.git" --no-deps !pip install -q opencv-python matplotlib scikit-learn scikit-image @@ -36,13 +66,13 @@ Install requirements ERROR: pip's dependency resolver does not currently take into account all the packages that are installed. This behaviour is the source of the following dependency conflicts. paddleclas 2.5.1 requires faiss-cpu==1.7.1.post2, but you have faiss-cpu 1.7.4 which is incompatible. paddleclas 2.5.1 requires gast==0.3.3, but you have gast 0.4.0 which is incompatible. - ppgan 2.1.0 requires librosa==0.8.1, but you have librosa 0.10.0.post2 which is incompatible. - ppgan 2.1.0 requires opencv-python<=4.6.0.66, but you have opencv-python 4.8.0.74 which is incompatible. + ppgan 2.1.0 requires librosa==0.8.1, but you have librosa 0.10.1 which is incompatible. + ppgan 2.1.0 requires opencv-python<=4.6.0.66, but you have opencv-python 4.8.0.76 which is incompatible. scikit-image 0.21.0 requires imageio>=2.27, but you have imageio 2.9.0 which is incompatible. -Imports -~~~~~~~ +Imports `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ .. code:: ipython3 @@ -85,8 +115,8 @@ Imports ) raise -Settings -~~~~~~~~ +Settings `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ .. code:: ipython3 @@ -100,8 +130,8 @@ Settings ir_path = model_path.with_suffix(".xml") onnx_path = model_path.with_suffix(".onnx") -Functions -~~~~~~~~~ +Functions `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ .. code:: ipython3 @@ -115,8 +145,8 @@ Functions image = cv2.resize(image, (max_width, new_height)) return image -Inference on PaddleGAN Model ----------------------------- +Inference on PaddleGAN Model `⇑ <#top>`__ +############################################################################################################################### The PaddleGAN `documentation `__ @@ -133,7 +163,7 @@ source of the function. .. parsed-literal:: - [07/11 23:03:15] ppgan INFO: Found /opt/home/k8sworker/.cache/ppgan/animeganv2_hayao.pdparams + [08/17 16:13:48] ppgan INFO: Found /opt/home/k8sworker/.cache/ppgan/animeganv2_hayao.pdparams .. code:: ipython3 @@ -211,8 +241,8 @@ cell. The anime image was saved to output/coco_bricks_anime_pg.jpg -Show Inference Results on PaddleGAN model -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Show Inference Results on PaddleGAN model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ .. code:: ipython3 @@ -228,15 +258,15 @@ Show Inference Results on PaddleGAN model .. image:: 206-vision-paddlegan-anime-with-output_files/206-vision-paddlegan-anime-with-output_15_0.png -Model Conversion to ONNX and OpenVINO IR ----------------------------------------- +Model Conversion to ONNX and OpenVINO IR `⇑ <#top>`__ +############################################################################################################################### Convert the PaddleGAN model to OpenVINO IR by first converting PaddleGAN to ONNX with ``paddle2onnx`` and then converting the ONNX model to -OpenVINO IR with Model Optimizer. +OpenVINO IR with model conversion API. -Convert to ONNX -~~~~~~~~~~~~~~~ +Convert to ONNX `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ Exporting to ONNX requires specifying an input shape with PaddlePaddle ``InputSpec`` and calling ``paddle.onnx.export``. Then, check the input @@ -268,23 +298,23 @@ succeeds, the output of the next cell will include .. parsed-literal:: - 2023-07-11 23:03:23 [INFO] Static PaddlePaddle model saved in model/paddle_model_static_onnx_temp_dir. + 2023-08-17 16:13:56 [INFO] Static PaddlePaddle model saved in model/paddle_model_static_onnx_temp_dir. [Paddle2ONNX] Start to parse PaddlePaddle model... [Paddle2ONNX] Model file path: model/paddle_model_static_onnx_temp_dir/model.pdmodel [Paddle2ONNX] Paramters file path: model/paddle_model_static_onnx_temp_dir/model.pdiparams [Paddle2ONNX] Start to parsing Paddle model... [Paddle2ONNX] Use opset_version = 11 for ONNX export. [Paddle2ONNX] PaddlePaddle model is exported as ONNX format now. - 2023-07-11 23:03:24 [INFO] ONNX model saved in model/paddlegan_anime.onnx. + 2023-08-17 16:13:56 [INFO] ONNX model saved in model/paddlegan_anime.onnx. .. parsed-literal:: - I0711 23:03:23.947630 3455115 interpretercore.cc:267] New Executor is Running. + I0817 16:13:56.664121 2277406 interpretercore.cc:237] New Executor is Running. -Convert to OpenVINO IR -~~~~~~~~~~~~~~~~~~~~~~ +Convert to OpenVINO IR `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ The OpenVINO IR format enables storing the preprocessing normalization in the model file. It is then no longer necessary to normalize input @@ -319,19 +349,18 @@ normalize uses a mean and scale of ``[127.5, 127.5, 127.5]``. The ``ResizeToScale`` class is called with ``(256,256)`` as the argument for size. Further analysis shows that this is the minimum size to resize to. The ``ResizeToScale`` class transform resizes images to the size -specified in the ``ResizeToScale`` params, with width and height as +specified in the ``ResizeToScale`` parameters, with width and height as multiples of 32. Once the mean and standard deviation values, and the shape of the model -inputs are known, you can use Model Optimizer and convert the model to -OpenVINO IR with these values. Use ``FP16`` precision and set log level -to ``CRITICAL`` to ignore warnings that are irrelevant for this demo. -For information about setting the parameters, see the `Model Optimizer -Documentation `__ -. +inputs are known, you can use model conversion API and convert the model +to OpenVINO IR with these values. Use ``FP16`` precision and set log +level to ``CRITICAL`` to ignore warnings that are irrelevant for this +demo. For information about setting the parameters, see this +`page `__. -**Convert ONNX Model to OpenVINO IR with**\ `Model Optimizer Python -API `__ +**Convert ONNX Model to OpenVINO IR with**\ `Model Conversion Python +API `__ .. code:: ipython3 @@ -357,19 +386,20 @@ API `__ Exporting ONNX model to OpenVINO IR... This may take a few minutes. -Show Inference Results on OpenVINO IR and PaddleGAN Models ----------------------------------------------------------- +Show Inference Results on OpenVINO IR and PaddleGAN Models `⇑ <#top>`__ +############################################################################################################################### -If the output of Model Optimizer in the cell above showed *SUCCESS*, the -model conversion succeeded and the OpenVINO IR model has been generated. +If the conversion is successful, the output of model conversion API in +the cell above will show *SUCCESS*, and the OpenVINO IR model will be +generated. Now, use the model for inference with the ``adjust_brightness()`` method from the PaddleGAN model. However, in order to use the OpenVINO IR model without installing PaddleGAN, it is useful to check what these functions do and extract them. -Create Postprocessing Functions -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Create Postprocessing Functions `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ .. code:: ipython3 @@ -412,8 +442,8 @@ OpenVINO IR model dstf = np.uint8(dstf) return dstf -Do Inference on OpenVINO IR Model -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Do Inference on OpenVINO IR Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ Load the OpenVINO IR model and do inference, following the same steps as for the PaddleGAN model. For more information about inference on @@ -424,12 +454,40 @@ The OpenVINO IR model is generated with an input shape that is computed based on the input image. If you do inference on images with different input shapes, results may differ from the PaddleGAN results. +Select inference device `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + core = Core() + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + .. code:: ipython3 # Load and prepare the IR model. - ie = Core() - model = ie.read_model(model=ir_path) - compiled_model = ie.compile_model(model=model, device_name="CPU") + core = Core() + model = core.read_model(model=ir_path) + compiled_model = core.compile_model(model=model, device_name=device.value) input_key = compiled_model.input(0) output_key = compiled_model.output(0) @@ -480,11 +538,12 @@ input shapes, results may differ from the PaddleGAN results. -.. image:: 206-vision-paddlegan-anime-with-output_files/206-vision-paddlegan-anime-with-output_36_0.png +.. image:: 206-vision-paddlegan-anime-with-output_files/206-vision-paddlegan-anime-with-output_37_0.png -Performance Comparison ----------------------- +Performance Comparison `⇑ <#top>`__ +############################################################################################################################### + Measure the time it takes to do inference on an image. This gives an indication of performance. It is not a perfect measure. Since the @@ -505,24 +564,6 @@ measure inference on one image. For more accurate benchmarking, use f"seconds per image, FPS: {NUM_IMAGES/time_ir:.2f}" ) - ## Uncomment the lines below to measure inference time on an Intel iGPU. - ## Note that it will take some time to load the model to GPU. - - # If "GPU" in ie.available_devices: - # # Loading the IR model on GPU takes some time. - # compiled_model = ie.compile_model(model=model, device_name="GPU") - # start = time.perf_counter() - # for _ in range(NUM_IMAGES): - # exec_net_multi([input_image]) - # end = time.perf_counter() - # time_ir = end - start - # print( - # f"OpenVINO IR model in OpenVINO Runtime/GPU: {time_ir/NUM_IMAGES:.3f} " - # f"seconds per image, FPS: {NUM_IMAGES/time_ir:.2f}" - # ) - # else: - # print("A supported iGPU device is not available on this system.") - ## `PADDLEGAN_INFERENCE` is defined in the "Inference on PaddleGAN model" section above. ## Uncomment the next line to enable a performance comparison with the PaddleGAN model ## if you disabled it earlier. @@ -544,19 +585,17 @@ measure inference on one image. For more accurate benchmarking, use .. parsed-literal:: - OpenVINO IR model in OpenVINO Runtime/CPU: 0.473 seconds per image, FPS: 2.12 - PaddleGAN model on CPU: 6.541 seconds per image, FPS: 0.15 + OpenVINO IR model in OpenVINO Runtime/CPU: 0.469 seconds per image, FPS: 2.13 + PaddleGAN model on CPU: 6.121 seconds per image, FPS: 0.16 -References ----------- +References `⇑ <#top>`__ +############################################################################################################################### - `PaddleGAN `__ - `Paddle2ONNX `__ -- `OpenVINO ONNX - support `__ -- `OpenVINO Model Optimizer - Documentation `__ +- `OpenVINO ONNX support `__ +- `Model Conversion API `__ The PaddleGAN code that is shown in this notebook is written by PaddlePaddle Authors and licensed under the Apache 2.0 license. The diff --git a/docs/notebooks/206-vision-paddlegan-anime-with-output_files/206-vision-paddlegan-anime-with-output_36_0.png b/docs/notebooks/206-vision-paddlegan-anime-with-output_files/206-vision-paddlegan-anime-with-output_37_0.png similarity index 100% rename from docs/notebooks/206-vision-paddlegan-anime-with-output_files/206-vision-paddlegan-anime-with-output_36_0.png rename to docs/notebooks/206-vision-paddlegan-anime-with-output_files/206-vision-paddlegan-anime-with-output_37_0.png diff --git a/docs/notebooks/206-vision-paddlegan-anime-with-output_files/index.html b/docs/notebooks/206-vision-paddlegan-anime-with-output_files/index.html index de060639c9c..c060ec41c78 100644 --- a/docs/notebooks/206-vision-paddlegan-anime-with-output_files/index.html +++ b/docs/notebooks/206-vision-paddlegan-anime-with-output_files/index.html @@ -1,8 +1,8 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/206-vision-paddlegan-anime-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/206-vision-paddlegan-anime-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/206-vision-paddlegan-anime-with-output_files/


../
-206-vision-paddlegan-anime-with-output_15_0.png    12-Jul-2023 00:11             1810982
-206-vision-paddlegan-anime-with-output_36_0.png    12-Jul-2023 00:11             1931653
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/206-vision-paddlegan-anime-with-output_files/


../
+206-vision-paddlegan-anime-with-output_15_0.png    16-Aug-2023 01:31             1810982
+206-vision-paddlegan-anime-with-output_37_0.png    16-Aug-2023 01:31             1931653
 

diff --git a/docs/notebooks/207-vision-paddlegan-superresolution-with-output.rst b/docs/notebooks/207-vision-paddlegan-superresolution-with-output.rst index 296fc664dc6..5967a0bf7b1 100644 --- a/docs/notebooks/207-vision-paddlegan-superresolution-with-output.rst +++ b/docs/notebooks/207-vision-paddlegan-superresolution-with-output.rst @@ -1,6 +1,8 @@ Super Resolution with PaddleGAN and OpenVINO™ ============================================= +.. _top: + This notebook demonstrates converting the RealSR (real-world super-resolution) model from `PaddlePaddle/PaddleGAN `__ @@ -16,8 +18,29 @@ from CVPR 2020. This notebook works best with small images (up to 800x600 resolution). -Imports -------- +**Table of contents**: + +- `Imports <#imports>`__ +- `Settings <#settings>`__ +- `Inference on PaddlePaddle Model <#inference-on-paddlepaddle-model>`__ + + - `Investigate PaddleGAN Model <#investigate-paddlegan-model>`__ + - `Do Inference <#do-inference>`__ + +- `Convert PaddleGAN Model to ONNX and OpenVINO IR <#convert-paddlegan-model-to-onnx-and-openvino-ir>`__ + + - `Convert PaddlePaddle Model to ONNX <#convert-paddlepaddle-model-to-onnx>`__ + - `Convert ONNX Model to OpenVINO IR with Model Conversion Python API <#convert-onnx-model-to-openvino-ir-with-model-conversion-python-api>`__ + +- `Do Inference on OpenVINO IR Model <#do-inference-on-openvino-ir-model>`__ + + - `Select inference device <#select-inference-device>`__ + - `Show an Animated GIF <#show-an-animated-gif>`__ + - `Create a Comparison Video <#create-a-comparison-video>`__ + +Imports `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -49,8 +72,9 @@ Imports sys.path.append("../utils") from notebook_utils import NotebookAlert -Settings --------- +Settings `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -64,11 +88,13 @@ Settings ir_path = model_path.with_suffix(".xml") onnx_path = model_path.with_suffix(".onnx") -Inference on PaddlePaddle Model -------------------------------- +Inference on PaddlePaddle Model `⇑ <#top>`__ +############################################################################################################################### + + +Investigate PaddleGAN Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Investigate PaddleGAN Model -~~~~~~~~~~~~~~~~~~~~~~~~~~~ The `PaddleGAN documentation `__ explains @@ -86,7 +112,7 @@ source code. .. parsed-literal:: - [07/11 23:03:50] ppgan INFO: Found /opt/home/k8sworker/.cache/ppgan/DF2K_JPEG.pdparams + [08/15 23:08:25] ppgan INFO: Found /opt/home/k8sworker/.cache/ppgan/DF2K_JPEG.pdparams .. code:: ipython3 @@ -124,8 +150,9 @@ To get more information about how the model looks like, use the # sr.model?? -Do Inference -~~~~~~~~~~~~ +Do Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + To show inference on the PaddlePaddle model, set ``PADDLEGAN_INFERENCE`` to ``True`` in the cell below. Keep in mind that performing inference @@ -169,15 +196,17 @@ may take some time. print(f"Inference duration: {duration:.2f} seconds") plt.imshow(result_image); -Convert PaddleGAN Model to ONNX and OpenVINO IR ------------------------------------------------ +Convert PaddleGAN Model to ONNX and OpenVINO IR `⇑ <#top>`__ +############################################################################################################################### + To convert the PaddlePaddle model to OpenVINO IR, first convert the model to ONNX, and then convert the ONNX model to the OpenVINO IR format. -Convert PaddlePaddle Model to ONNX -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Convert PaddlePaddle Model to ONNX `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -194,12 +223,12 @@ Convert PaddlePaddle Model to ONNX .. parsed-literal:: - 2023-07-11 23:03:56 [INFO] Static PaddlePaddle model saved in model/paddle_model_static_onnx_temp_dir. + 2023-08-15 23:08:31 [INFO] Static PaddlePaddle model saved in model/paddle_model_static_onnx_temp_dir. .. parsed-literal:: - I0711 23:03:56.497538 3455488 interpretercore.cc:267] New Executor is Running. + I0815 23:08:31.864205 2062621 interpretercore.cc:267] New Executor is Running. .. parsed-literal:: @@ -210,18 +239,18 @@ Convert PaddlePaddle Model to ONNX [Paddle2ONNX] Start to parsing Paddle model... [Paddle2ONNX] Use opset_version = 13 for ONNX export. [Paddle2ONNX] PaddlePaddle model is exported as ONNX format now. - 2023-07-11 23:04:00 [INFO] ONNX model saved in model/paddlegan_sr.onnx. + 2023-08-15 23:08:35 [INFO] ONNX model saved in model/paddlegan_sr.onnx. -Convert ONNX Model to OpenVINO IR with `Model Optimizer Python API `__ -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Convert ONNX Model to OpenVINO IR with Model Conversion Python API `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ .. code:: ipython3 from openvino.tools import mo from openvino.runtime import serialize - ## Uncomment the command below to show Model Optimizer help, which shows the possible arguments for Model Optimizer. + ## Uncomment the command below to show help, which shows the possible arguments for model conversion API. # mo.convert_model(help=True) .. code:: ipython3 @@ -243,17 +272,46 @@ Convert ONNX Model to OpenVINO IR with `Model Optimizer Python API `__ +############################################################################################################################### + .. code:: ipython3 # Read the network and get input and output names. - ie = Core() + core = Core() # Alternatively, the model obtained from `mo.convert_model()` may be used here - model = ie.read_model(model=ir_path) + model = core.read_model(model=ir_path) input_layer = model.input(0) +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + .. code:: ipython3 # Load and show the image. @@ -273,18 +331,18 @@ Do Inference on OpenVINO IR Model .. parsed-literal:: - + -.. image:: 207-vision-paddlegan-superresolution-with-output_files/207-vision-paddlegan-superresolution-with-output_25_1.png +.. image:: 207-vision-paddlegan-superresolution-with-output_files/207-vision-paddlegan-superresolution-with-output_27_1.png .. code:: ipython3 # Load the network to the CPU device (this may take a few seconds). - compiled_model = ie.compile_model(model=model, device_name="CPU") + compiled_model = core.compile_model(model=model, device_name=device.value) output_layer = compiled_model.output(0) .. code:: ipython3 @@ -302,7 +360,7 @@ Do Inference on OpenVINO IR Model .. parsed-literal:: - Inference duration: 3.26 seconds + Inference duration: 3.30 seconds .. code:: ipython3 @@ -325,16 +383,17 @@ Do Inference on OpenVINO IR Model .. parsed-literal:: - + -.. image:: 207-vision-paddlegan-superresolution-with-output_files/207-vision-paddlegan-superresolution-with-output_29_1.png +.. image:: 207-vision-paddlegan-superresolution-with-output_files/207-vision-paddlegan-superresolution-with-output_31_1.png -Show an Animated GIF -~~~~~~~~~~~~~~~~~~~~ +Show an Animated GIF `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + To visualize the difference between the bicubic image and the superresolution image, create an animated GIF image that switches @@ -362,13 +421,14 @@ between both versions. -.. image:: 207-vision-paddlegan-superresolution-with-output_files/207-vision-paddlegan-superresolution-with-output_31_0.png +.. image:: 207-vision-paddlegan-superresolution-with-output_files/207-vision-paddlegan-superresolution-with-output_33_0.png :width: 960px -Create a Comparison Video -~~~~~~~~~~~~~~~~~~~~~~~~~ +Create a Comparison Video `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Create a video with a “slider”, showing the bicubic image to the right and the superresolution image on the left. diff --git a/docs/notebooks/207-vision-paddlegan-superresolution-with-output_files/207-vision-paddlegan-superresolution-with-output_25_1.png b/docs/notebooks/207-vision-paddlegan-superresolution-with-output_files/207-vision-paddlegan-superresolution-with-output_27_1.png similarity index 100% rename from docs/notebooks/207-vision-paddlegan-superresolution-with-output_files/207-vision-paddlegan-superresolution-with-output_25_1.png rename to docs/notebooks/207-vision-paddlegan-superresolution-with-output_files/207-vision-paddlegan-superresolution-with-output_27_1.png diff --git a/docs/notebooks/207-vision-paddlegan-superresolution-with-output_files/207-vision-paddlegan-superresolution-with-output_29_1.png b/docs/notebooks/207-vision-paddlegan-superresolution-with-output_files/207-vision-paddlegan-superresolution-with-output_31_1.png similarity index 100% rename from docs/notebooks/207-vision-paddlegan-superresolution-with-output_files/207-vision-paddlegan-superresolution-with-output_29_1.png rename to docs/notebooks/207-vision-paddlegan-superresolution-with-output_files/207-vision-paddlegan-superresolution-with-output_31_1.png diff --git a/docs/notebooks/207-vision-paddlegan-superresolution-with-output_files/207-vision-paddlegan-superresolution-with-output_31_0.png b/docs/notebooks/207-vision-paddlegan-superresolution-with-output_files/207-vision-paddlegan-superresolution-with-output_33_0.png similarity index 100% rename from docs/notebooks/207-vision-paddlegan-superresolution-with-output_files/207-vision-paddlegan-superresolution-with-output_31_0.png rename to docs/notebooks/207-vision-paddlegan-superresolution-with-output_files/207-vision-paddlegan-superresolution-with-output_33_0.png diff --git a/docs/notebooks/207-vision-paddlegan-superresolution-with-output_files/index.html b/docs/notebooks/207-vision-paddlegan-superresolution-with-output_files/index.html index 38f1acb8183..27b987aa18c 100644 --- a/docs/notebooks/207-vision-paddlegan-superresolution-with-output_files/index.html +++ b/docs/notebooks/207-vision-paddlegan-superresolution-with-output_files/index.html @@ -1,9 +1,9 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/207-vision-paddlegan-superresolution-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/207-vision-paddlegan-superresolution-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/207-vision-paddlegan-superresolution-with-output_files/


../
-207-vision-paddlegan-superresolution-with-outpu..> 12-Jul-2023 00:11              436999
-207-vision-paddlegan-superresolution-with-outpu..> 12-Jul-2023 00:11              476190
-207-vision-paddlegan-superresolution-with-outpu..> 12-Jul-2023 00:11             2835354
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/207-vision-paddlegan-superresolution-with-output_files/


../
+207-vision-paddlegan-superresolution-with-outpu..> 16-Aug-2023 01:31              436999
+207-vision-paddlegan-superresolution-with-outpu..> 16-Aug-2023 01:31              476190
+207-vision-paddlegan-superresolution-with-outpu..> 16-Aug-2023 01:31             2835354
 

diff --git a/docs/notebooks/208-optical-character-recognition-with-output.rst b/docs/notebooks/208-optical-character-recognition-with-output.rst index bd3ef8b949c..79845f408ee 100644 --- a/docs/notebooks/208-optical-character-recognition-with-output.rst +++ b/docs/notebooks/208-optical-character-recognition-with-output.rst @@ -1,6 +1,8 @@ Optical Character Recognition (OCR) with OpenVINO™ ================================================== +.. _top: + This tutorial demonstrates how to perform optical character recognition (OCR) with OpenVINO models. It is a continuation of the `004-hello-detection <004-hello-detection-with-output.html>`__ @@ -19,8 +21,34 @@ Zoo `__. For more information, refer to the `104-model-tools <104-model-tools-with-output.html>`__ tutorial. -Imports -------- +**Table of contents**: + +- `Imports <#imports>`__ +- `Settings <#settings>`__ +- `Download Models <#download-models>`__ +- `Convert Models <#convert-models>`__ +- `Select inference device <#select-inference-device>`__ +- `Object Detection <#object-detection>`__ + + - `Load a Detection Model <#load-a-detection-model>`__ + - `Load an Image <#load-an-image>`__ + - `Do Inference <#do-inference>`__ + - `Get Detection Results <#get-detection-results>`__ + +- `Text Recognition <#text-recognition>`__ + + - `Load Text Recognition Model <#load-text-recognition-model>`__ + - `Do Inference <#do-inference>`__ + +- `Show Results <#show-results>`__ + + - `Show Detected Text Boxes and OCR Results for the Image <#show-detected-text-boxes-and-ocr-results-for-the-image>`__ + - `Show the OCR Result per Bounding Box <#show-the-ocr-result-per-bounding-box>`__ + - `Print Annotations in Plain Text Format <#print-annotations-in-plain-text-format>`__ + +Imports `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -37,12 +65,13 @@ Imports sys.path.append("../utils") from notebook_utils import load_image -Settings --------- +Settings `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 - ie = Core() + core = Core() model_dir = Path("model") precision = "FP16" @@ -51,8 +80,9 @@ Settings model_dir.mkdir(exist_ok=True) -Download Models ---------------- +Download Models `⇑ <#top>`__ +############################################################################################################################### + The next cells will run Model Downloader to download the detection and recognition models. If the models have been downloaded before, they will @@ -250,6 +280,8 @@ Downloading horizontal-text-detection-0001, text-recognition-resnet-fc… ========== Replacing text in model/public/text-recognition-resnet-fc/vedastr/utils/config.py ========== Replacing text in model/public/text-recognition-resnet-fc/vedastr/utils/config.py ========== Replacing text in model/public/text-recognition-resnet-fc/vedastr/utils/config.py + ========== Replacing text in model/public/text-recognition-resnet-fc/vedastr/models/bodies/feature_extractors/encoders/backbones/resnet.py + ========== Replacing text in model/public/text-recognition-resnet-fc/vedastr/models/bodies/feature_extractors/encoders/backbones/resnet.py ========== Unpacking model/public/text-recognition-resnet-fc/vedastr/addict-2.4.0-py3-none-any.whl @@ -267,8 +299,9 @@ text-recognition-resnet-fc. # for line in download_result: # print(line) -Convert Models --------------- +Convert Models `⇑ <#top>`__ +############################################################################################################################### + The downloaded detection model is an Intel model, which is already in OpenVINO Intermediate Representation (OpenVINO IR) format. The text @@ -300,45 +333,74 @@ Converting text-recognition-resnet-fc… .. parsed-literal:: ========== Converting text-recognition-resnet-fc to ONNX - Conversion to ONNX command: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/bin/python -- /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/model_zoo/internal_scripts/pytorch_to_onnx.py --model-path=/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/model_zoo/models/public/text-recognition-resnet-fc --model-path=model/public/text-recognition-resnet-fc --model-name=get_model --import-module=model '--model-param=file_config=r"model/public/text-recognition-resnet-fc/vedastr/configs/resnet_fc.py"' '--model-param=weights=r"model/public/text-recognition-resnet-fc/vedastr/ckpt/resnet_fc.pth"' --input-shape=1,1,32,100 --input-names=input --output-names=output --output-file=model/public/text-recognition-resnet-fc/resnet_fc.onnx + Conversion to ONNX command: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/bin/python -- /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/model_zoo/internal_scripts/pytorch_to_onnx.py --model-path=/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/model_zoo/models/public/text-recognition-resnet-fc --model-path=model/public/text-recognition-resnet-fc --model-name=get_model --import-module=model '--model-param=file_config=r"model/public/text-recognition-resnet-fc/vedastr/configs/resnet_fc.py"' '--model-param=weights=r"model/public/text-recognition-resnet-fc/vedastr/ckpt/resnet_fc.pth"' --input-shape=1,1,32,100 --input-names=input --output-names=output --output-file=model/public/text-recognition-resnet-fc/resnet_fc.onnx - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torchvision/models/_utils.py:252: UserWarning: Accessing the model URLs via the internal dictionary of the module is deprecated since 0.13 and may be removed in the future. Please access them via the appropriate Weights Enum instead. - warnings.warn( ONNX check passed successfully. ========== Converting text-recognition-resnet-fc to IR (FP16) - Conversion command: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/bin/python -- /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/bin/mo --framework=onnx --output_dir=/tmp/tmpvegltj_7 --model_name=text-recognition-resnet-fc --input=input '--mean_values=input[127.5]' '--scale_values=input[127.5]' --output=output --input_model=model/public/text-recognition-resnet-fc/resnet_fc.onnx '--layout=input(NCHW)' '--input_shape=[1, 1, 32, 100]' --compress_to_fp16=True + Conversion command: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/bin/python -- /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/bin/mo --framework=onnx --output_dir=/tmp/tmppkwl27u7 --model_name=text-recognition-resnet-fc --input=input '--mean_values=input[127.5]' '--scale_values=input[127.5]' --output=output --input_model=model/public/text-recognition-resnet-fc/resnet_fc.onnx '--layout=input(NCHW)' '--input_shape=[1, 1, 32, 100]' --compress_to_fp16=True [ INFO ] Generated IR will be compressed to FP16. If you get lower accuracy, please consider disabling compression by removing argument --compress_to_fp16 or set it to false --compress_to_fp16=False. - Find more information about compression to FP16 at https://docs.openvino.ai/latest/openvino_docs_MO_DG_FP16_Compression.html + Find more information about compression to FP16 at https://docs.openvino.ai/2023.0/openvino_docs_MO_DG_FP16_Compression.html [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. - Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/latest/openvino_2_0_transition_guide.html + Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html [ SUCCESS ] Generated IR version 11 model. - [ SUCCESS ] XML file: /tmp/tmpvegltj_7/text-recognition-resnet-fc.xml - [ SUCCESS ] BIN file: /tmp/tmpvegltj_7/text-recognition-resnet-fc.bin + [ SUCCESS ] XML file: /tmp/tmppkwl27u7/text-recognition-resnet-fc.xml + [ SUCCESS ] BIN file: /tmp/tmppkwl27u7/text-recognition-resnet-fc.bin -Object Detection ----------------- +Select inference device `⇑ <#top>`__ +############################################################################################################################### + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +Object Detection `⇑ <#top>`__ +############################################################################################################################### + Load a detection model, load an image, do inference and get the detection inference result. -Load a Detection Model -~~~~~~~~~~~~~~~~~~~~~~ +Load a Detection Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 - detection_model = ie.read_model( + detection_model = core.read_model( model=detection_model_path, weights=detection_model_path.with_suffix(".bin") ) - detection_compiled_model = ie.compile_model(model=detection_model, device_name="CPU") + detection_compiled_model = core.compile_model(model=detection_model, device_name=device.value) detection_input_layer = detection_compiled_model.input(0) -Load an Image -~~~~~~~~~~~~~ +Load an Image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -360,11 +422,12 @@ Load an Image -.. image:: 208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_13_0.png +.. image:: 208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_15_0.png -Do Inference -~~~~~~~~~~~~ +Do Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Text boxes are detected in the images and returned as blobs of data in the shape of ``[100, 5]``. Each description of detection has the @@ -378,8 +441,9 @@ the shape of ``[100, 5]``. Each description of detection has the # Remove zero only boxes. boxes = boxes[~np.all(boxes == 0, axis=1)] -Get Detection Results -~~~~~~~~~~~~~~~~~~~~~ +Get Detection Results `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -447,22 +511,24 @@ Get Detection Results return rgb_image -Text Recogntion ---------------- +Text Recognition `⇑ <#top>`__ +############################################################################################################################### + Load the text recognition model and do inference on the detected boxes from the detection model. -Load Text Recognition Model -~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Load Text Recognition Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 - recognition_model = ie.read_model( + recognition_model = core.read_model( model=recognition_model_path, weights=recognition_model_path.with_suffix(".bin") ) - recognition_compiled_model = ie.compile_model(model=recognition_model, device_name="CPU") + recognition_compiled_model = core.compile_model(model=recognition_model, device_name=device.value) recognition_output_layer = recognition_compiled_model.output(0) recognition_input_layer = recognition_compiled_model.input(0) @@ -470,8 +536,9 @@ Load Text Recognition Model # Get the height and width of the input layer. _, _, H, W = recognition_input_layer.shape -Do Inference -~~~~~~~~~~~~ +Do Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -516,11 +583,13 @@ Do Inference boxes_with_annotations = list(zip(boxes, annotations)) -Show Results ------------- +Show Results `⇑ <#top>`__ +############################################################################################################################### + + +Show Detected Text Boxes and OCR Results for the Image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Show Detected Text Boxes and OCR Results for the Image -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Visualize the result by drawing boxes around recognized text and showing the OCR result from the text recognition model. @@ -532,11 +601,12 @@ the OCR result from the text recognition model. -.. image:: 208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_23_0.png +.. image:: 208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_0.png -Show the OCR Result per Bounding Box -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Show the OCR Result per Bounding Box `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Depending on the image, the OCR result may not be readable in the image with boxes, as displayed in the cell above. Use the code below to @@ -549,7 +619,7 @@ display the extracted boxes and the OCR result per box. -.. image:: 208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_0.png +.. image:: 208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_0.png @@ -557,7 +627,7 @@ building -.. image:: 208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_2.png +.. image:: 208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_2.png @@ -565,7 +635,7 @@ noyce -.. image:: 208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_4.png +.. image:: 208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_4.png @@ -573,7 +643,7 @@ noyce -.. image:: 208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_6.png +.. image:: 208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_6.png @@ -581,7 +651,7 @@ n -.. image:: 208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_8.png +.. image:: 208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_8.png @@ -589,15 +659,16 @@ center -.. image:: 208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_10.png +.. image:: 208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_10.png robert -Print Annotations in Plain Text Format -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Print Annotations in Plain Text Format `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Print annotations for detected text based on their position in the input image, starting from the upper left corner. diff --git a/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_13_0.png b/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_15_0.png similarity index 100% rename from docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_13_0.png rename to docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_15_0.png diff --git a/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_23_0.png b/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_23_0.png deleted file mode 100644 index dbb9a059a92..00000000000 --- a/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_23_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:3cf5c39f43e8ae8b58bb2cf83ebb5d94043a715d59e1f1a5641a83742d17b7a4 -size 923631 diff --git a/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_0.png b/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_0.png index bd61b72ac14..dbb9a059a92 100644 --- a/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_0.png +++ b/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_0.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:3afa4268972711c6dd52892c4fa0f9600df3ae40bbec42d5867b83fd75f9308c -size 11367 +oid sha256:3cf5c39f43e8ae8b58bb2cf83ebb5d94043a715d59e1f1a5641a83742d17b7a4 +size 923631 diff --git a/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_0.jpg b/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_0.jpg similarity index 100% rename from docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_0.jpg rename to docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_0.jpg diff --git a/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_0.png b/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_0.png new file mode 100644 index 00000000000..bd61b72ac14 --- /dev/null +++ b/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:3afa4268972711c6dd52892c4fa0f9600df3ae40bbec42d5867b83fd75f9308c +size 11367 diff --git a/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_10.jpg b/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_10.jpg similarity index 100% rename from docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_10.jpg rename to docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_10.jpg diff --git a/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_10.png b/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_10.png similarity index 100% rename from docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_10.png rename to docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_10.png diff --git a/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_2.jpg b/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_2.jpg similarity index 100% rename from docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_2.jpg rename to docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_2.jpg diff --git a/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_2.png b/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_2.png similarity index 100% rename from docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_2.png rename to docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_2.png diff --git a/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_4.jpg b/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_4.jpg similarity index 100% rename from docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_4.jpg rename to docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_4.jpg diff --git a/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_4.png b/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_4.png similarity index 100% rename from docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_4.png rename to docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_4.png diff --git a/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_6.jpg b/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_6.jpg similarity index 100% rename from docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_6.jpg rename to docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_6.jpg diff --git a/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_6.png b/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_6.png similarity index 100% rename from docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_6.png rename to docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_6.png diff --git a/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_8.jpg b/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_8.jpg similarity index 100% rename from docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_8.jpg rename to docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_8.jpg diff --git a/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_8.png b/docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_8.png similarity index 100% rename from docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_25_8.png rename to docs/notebooks/208-optical-character-recognition-with-output_files/208-optical-character-recognition-with-output_27_8.png diff --git a/docs/notebooks/208-optical-character-recognition-with-output_files/index.html b/docs/notebooks/208-optical-character-recognition-with-output_files/index.html index bf60dbf483b..09bf331cf60 100644 --- a/docs/notebooks/208-optical-character-recognition-with-output_files/index.html +++ b/docs/notebooks/208-optical-character-recognition-with-output_files/index.html @@ -1,20 +1,20 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/208-optical-character-recognition-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/208-optical-character-recognition-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/208-optical-character-recognition-with-output_files/


../
-208-optical-character-recognition-with-output_1..> 12-Jul-2023 00:11              305482
-208-optical-character-recognition-with-output_2..> 12-Jul-2023 00:11              923631
-208-optical-character-recognition-with-output_2..> 12-Jul-2023 00:11                1996
-208-optical-character-recognition-with-output_2..> 12-Jul-2023 00:11               11367
-208-optical-character-recognition-with-output_2..> 12-Jul-2023 00:11                1990
-208-optical-character-recognition-with-output_2..> 12-Jul-2023 00:11               11142
-208-optical-character-recognition-with-output_2..> 12-Jul-2023 00:11                1630
-208-optical-character-recognition-with-output_2..> 12-Jul-2023 00:11                8428
-208-optical-character-recognition-with-output_2..> 12-Jul-2023 00:11                 949
-208-optical-character-recognition-with-output_2..> 12-Jul-2023 00:11                2274
-208-optical-character-recognition-with-output_2..> 12-Jul-2023 00:11                 817
-208-optical-character-recognition-with-output_2..> 12-Jul-2023 00:11                1559
-208-optical-character-recognition-with-output_2..> 12-Jul-2023 00:11                 838
-208-optical-character-recognition-with-output_2..> 12-Jul-2023 00:11                1487
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/208-optical-character-recognition-with-output_files/


../
+208-optical-character-recognition-with-output_1..> 16-Aug-2023 01:31              305482
+208-optical-character-recognition-with-output_2..> 16-Aug-2023 01:31              923631
+208-optical-character-recognition-with-output_2..> 16-Aug-2023 01:31                1996
+208-optical-character-recognition-with-output_2..> 16-Aug-2023 01:31               11367
+208-optical-character-recognition-with-output_2..> 16-Aug-2023 01:31                1990
+208-optical-character-recognition-with-output_2..> 16-Aug-2023 01:31               11142
+208-optical-character-recognition-with-output_2..> 16-Aug-2023 01:31                1630
+208-optical-character-recognition-with-output_2..> 16-Aug-2023 01:31                8428
+208-optical-character-recognition-with-output_2..> 16-Aug-2023 01:31                 949
+208-optical-character-recognition-with-output_2..> 16-Aug-2023 01:31                2274
+208-optical-character-recognition-with-output_2..> 16-Aug-2023 01:31                 817
+208-optical-character-recognition-with-output_2..> 16-Aug-2023 01:31                1559
+208-optical-character-recognition-with-output_2..> 16-Aug-2023 01:31                 838
+208-optical-character-recognition-with-output_2..> 16-Aug-2023 01:31                1487
 

diff --git a/docs/notebooks/209-handwritten-ocr-with-output.rst b/docs/notebooks/209-handwritten-ocr-with-output.rst index a05e00ca061..e0f5913988f 100644 --- a/docs/notebooks/209-handwritten-ocr-with-output.rst +++ b/docs/notebooks/209-handwritten-ocr-with-output.rst @@ -1,6 +1,8 @@ Handwritten Chinese and Japanese OCR with OpenVINO™ =================================================== +.. _top: + In this tutorial, we perform optical character recognition (OCR) for handwritten Chinese (simplified) and Japanese. An OCR tutorial using the Latin alphabet is available in `notebook @@ -8,18 +10,34 @@ Latin alphabet is available in `notebook This model is capable of processing only one line of symbols at a time. The models used in this notebook are -`handwritten-japanese-recognition `__ +`handwritten-japanese-recognition-0001 `__ and -`handwritten-simplified-chinese `__. +`handwritten-simplified-chinese-0001 `__. To decode model outputs as readable text `kondate_nakayosi `__ and `scut_ept `__ -charlists are used. Both models are available on `Open Model -Zoo `__. +charlists are used. Both models are available on `Open Model Zoo `__. + +**Table of contents**: + +- `Imports <#imports>`__ +- `Settings <#settings>`__ +- `Select a Language <#select-a-language>`__ +- `Download the Model <#download-the-model>`__ +- `Load the Model and Execute <#load-the-model-and-execute>`__ +- `Select inference device <#select-inference-device>`__ +- `Fetch Information About Input and Output Layers <#fetch-information-about-input-and-output-layers>`__ +- `Load an Image <#load-an-image>`__ +- `Visualize Input Image <#visualize-input-image>`__ +- `Prepare Charlist <#prepare-charlist>`__ +- `Run Inference <#run-inference>`__ +- `Process the Output Data <#process-the-output-data>`__ +- `Print the Output <#print-the-output>`__ + +Imports `⇑ <#top>`__ +############################################################################################################################### -Imports -------- .. code:: ipython3 @@ -32,8 +50,9 @@ Imports import numpy as np from openvino.runtime import Core -Settings --------- +Settings `⇑ <#top>`__ +############################################################################################################################### + Set up all constants and folders used in this notebook @@ -66,8 +85,9 @@ To group files, you have to define the collection. In this case, use demo_image_name="handwritten_japanese_test.png", ) -Select a Language ------------------ +Select a Language `⇑ <#top>`__ +############################################################################################################################### + Depending on your choice you will need to change a line of code in the cell below. @@ -84,8 +104,9 @@ If you want to perform OCR on a text in Japanese, set selected_language = languages.get(language) -Download the Model ------------------- +Download the Model `⇑ <#top>`__ +############################################################################################################################### + In addition to images and charlists, you need to download the model file. In the sections below, there are cells for downloading either the @@ -120,8 +141,9 @@ and downloads the selected model. -Load the Network and Execute ----------------------------- +Load the Model and Execute `⇑ <#top>`__ +############################################################################################################################### + When all files are downloaded and language is selected, read and compile the network to run inference. The path to the model is defined based on @@ -129,29 +151,47 @@ the selected language. .. code:: ipython3 - ie = Core() + core = Core() path_to_model = path_to_model_weights.with_suffix(".xml") - model = ie.read_model(model=path_to_model) + model = core.read_model(model=path_to_model) -Select a Device Name -~~~~~~~~~~~~~~~~~~~~ +Select inference device `⇑ <#top>`__ +############################################################################################################################### -You may choose to run the network on multiple devices. By default, it -will load the model on CPU (CPU, GPU, etc. can be set manually) or let -the engine choose the best available device (AUTO). -To list all available devices, uncomment and run the line -``print(ie.available_devices)``. +Select device from dropdown list for running inference using OpenVINO: .. code:: ipython3 - # To check available device names run the line below - # print(ie.available_devices) + import ipywidgets as widgets - compiled_model = ie.compile_model(model=model, device_name="CPU") + core = Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + compiled_model = core.compile_model(model=model, device_name=device.value) + +Fetch Information About Input and Output Layers `⇑ <#top>`__ +############################################################################################################################### -Fetch Information About Input and Output Layers ------------------------------------------------ Now that the model is loaded, fetch information about the input and output layers (shape). @@ -161,8 +201,9 @@ output layers (shape). recognition_output_layer = compiled_model.output(0) recognition_input_layer = compiled_model.input(0) -Load an Image -------------- +Load an Image `⇑ <#top>`__ +############################################################################################################################### + Next, load an image. The model expects a single-channel image as input, so the image is read in grayscale. @@ -206,8 +247,9 @@ keep letters proportional and meet input shape. # Reshape to network input shape. input_image = resized_image[None, None, :, :] -Visualise Input Image ---------------------- +Visualize Input Image `⇑ <#top>`__ +############################################################################################################################### + After preprocessing, you can display the image. @@ -219,11 +261,12 @@ After preprocessing, you can display the image. -.. image:: 209-handwritten-ocr-with-output_files/209-handwritten-ocr-with-output_20_0.png +.. image:: 209-handwritten-ocr-with-output_files/209-handwritten-ocr-with-output_21_0.png -Prepare Charlist ----------------- +Prepare Charlist `⇑ <#top>`__ +############################################################################################################################### + The model is loaded and the image is ready. The only element left is the charlist, which is downloaded. You must add a blank symbol at the @@ -241,8 +284,9 @@ Chinese and Japanese models. with open(f"{charlist_folder}/{used_charlist}", "r", encoding="utf-8") as charlist: letters = blank_char + "".join(line.strip() for line in charlist) -Run Inference -------------- +Run Inference `⇑ <#top>`__ +############################################################################################################################### + Now, run inference. The ``compiled_model()`` function takes a list with input(s) in the same order as model input(s). Then, fetch the output @@ -253,8 +297,9 @@ from output tensors. # Run inference on the model predictions = compiled_model([input_image])[recognition_output_layer] -Process the Output Data ------------------------ +Process the Output Data `⇑ <#top>`__ +############################################################################################################################### + The output of a model is in the ``W x B x L`` format, where: @@ -293,8 +338,9 @@ Finally, get the symbols from corresponding indexes in the charlist. # Assign letters to indexes from the output array. output_text = [letters[letter_index] for letter_index in output_text_indexes] -Print the Output ----------------- +Print the Output `⇑ <#top>`__ +############################################################################################################################### + Now, having a list of letters predicted by the model, you can display the image with predicted text printed below. @@ -314,5 +360,5 @@ the image with predicted text printed below. -.. image:: 209-handwritten-ocr-with-output_files/209-handwritten-ocr-with-output_29_1.png +.. image:: 209-handwritten-ocr-with-output_files/209-handwritten-ocr-with-output_30_1.png diff --git a/docs/notebooks/209-handwritten-ocr-with-output_files/209-handwritten-ocr-with-output_20_0.png b/docs/notebooks/209-handwritten-ocr-with-output_files/209-handwritten-ocr-with-output_21_0.png similarity index 100% rename from docs/notebooks/209-handwritten-ocr-with-output_files/209-handwritten-ocr-with-output_20_0.png rename to docs/notebooks/209-handwritten-ocr-with-output_files/209-handwritten-ocr-with-output_21_0.png diff --git a/docs/notebooks/209-handwritten-ocr-with-output_files/209-handwritten-ocr-with-output_29_1.png b/docs/notebooks/209-handwritten-ocr-with-output_files/209-handwritten-ocr-with-output_30_1.png similarity index 100% rename from docs/notebooks/209-handwritten-ocr-with-output_files/209-handwritten-ocr-with-output_29_1.png rename to docs/notebooks/209-handwritten-ocr-with-output_files/209-handwritten-ocr-with-output_30_1.png diff --git a/docs/notebooks/209-handwritten-ocr-with-output_files/index.html b/docs/notebooks/209-handwritten-ocr-with-output_files/index.html index cf187cc4968..2e8667d2f9a 100644 --- a/docs/notebooks/209-handwritten-ocr-with-output_files/index.html +++ b/docs/notebooks/209-handwritten-ocr-with-output_files/index.html @@ -1,8 +1,8 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/209-handwritten-ocr-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/209-handwritten-ocr-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/209-handwritten-ocr-with-output_files/


../
-209-handwritten-ocr-with-output_20_0.png           12-Jul-2023 00:11               53571
-209-handwritten-ocr-with-output_29_1.png           12-Jul-2023 00:11               53571
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/209-handwritten-ocr-with-output_files/


../
+209-handwritten-ocr-with-output_21_0.png           16-Aug-2023 01:31               53571
+209-handwritten-ocr-with-output_30_1.png           16-Aug-2023 01:31               53571
 

diff --git a/docs/notebooks/210-slowfast-video-recognition-with-output.rst b/docs/notebooks/210-slowfast-video-recognition-with-output.rst index c4340ecee7c..e795d99a6ef 100644 --- a/docs/notebooks/210-slowfast-video-recognition-with-output.rst +++ b/docs/notebooks/210-slowfast-video-recognition-with-output.rst @@ -1,6 +1,8 @@ Video Recognition using SlowFast and OpenVINO™ ============================================== +.. _top: + Teaching machines to detect, understand and analyze the contents of images has been one of the more well-known and well-studied problems in computer vision. However, analyzing videos to understand what is @@ -18,7 +20,7 @@ ability to effectively capture both fast and slow-motion information in video sequences, making it particularly well-suited to tasks that require a temporal and spatial understanding of the data. -.. image:: https://camo.githubusercontent.com/8d7441ad72421909352b2c3010aae2ee3f4463ef8c14a0a0e1cb92cfb7f6f9f6/68747470733a2f2f757365722d696d616765732e67697468756275736572636f6e74656e742e636f6d2f33343332343135352f3134333034343131312d39343637366636342d376261382d343038312d393031312d6638303534626564373033302e706e67 +|image0| More details about the network can be found in the original `paper `__ @@ -36,18 +38,35 @@ This tutorial consists of the following steps - Convert ONNX Model to OpenVINO Intermediate Representation - Verify inference with the converted model -Prepare PyTorch Model ---------------------- +.. |image0| image:: https://user-images.githubusercontent.com/34324155/143044111-94676f64-7ba8-4081-9011-f8054bed7030.png + +**Table of contents**: + +- `Prepare PyTorch Model <#prepare-pytorch-model>`__ + + - `Install necessary packages <#install-necessary-packages>`__ + - `Imports and Settings <#imports-and-settings>`__ + +- `Export to ONNX <#export-to-onnx>`__ +- `Convert ONNX to OpenVINO™ Intermediate Representation <#convert-onnx-to-openvino™-intermediate-representation>`__ +- `Select inference device <#select-inference-device>`__ +- `Verify Model Inference <#verify-model-inference>`__ + +Prepare PyTorch Model `⇑ <#top>`__ +############################################################################################################################### + + +Install necessary packages `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Install necessary packages -~~~~~~~~~~~~~~~~~~~~~~~~~~ .. code:: ipython3 !pip install fvcore -q -Imports and Settings -~~~~~~~~~~~~~~~~~~~~ +Imports and Settings `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -899,13 +918,14 @@ inference using the same. The top 5 predictions can be seen below. Predicted labels: archery, throwing axe, playing paintball, golf driving, riding or walking with horse -Export to ONNX --------------- +Export to ONNX `⇑ <#top>`__ +############################################################################################################################### + Now that we have obtained our trained model and checked inference with it, we export the PyTorch model to Open Neural Network Exchange(ONNX) format, an open format for representing machine learning models, so that -we can use the Model Optimizer to convert it to OpenVINO Intermediate +we can use model conversion API to convert it to OpenVINO Intermediate Representation format(IR). This can be later used to run inference using the OpenVINO Runtime. Note that although the OpenVINO Runtime supports running ONNX models directly, converting to IR format enables us to take @@ -923,14 +943,15 @@ quantization. export_params=True, ) -Convert ONNX to OpenVINO™ Intermediate Representation ------------------------------------------------------ +Convert ONNX to OpenVINO™ Intermediate Representation `⇑ <#top>`__ +############################################################################################################################### + Now that our ONNX model is ready, we can convert it to IR format. In -this format, the network is represented using two files: an xml file +this format, the network is represented using two files: an ``xml`` file describing the network architecture and an accompanying binary file that stores constant values such as convolution weights in a binary format. -We can use the OpenVINO Model Optimizer for converting into IR format as +We can use model conversion API for converting into IR format as follows. The ``convert_model`` method returns an ``openvino.runtime.Model`` object that can either be compiled and inferred or serialized. @@ -960,12 +981,43 @@ using the ``weights`` parameter. # read converted model conv_model = core.read_model(str(IR_PATH)) - - # load model on CPU device - compiled_model = core.compile_model(model=conv_model, device_name="CPU") -Verify Model Inference ----------------------- +Select inference device `⇑ <#top>`__ +############################################################################################################################### + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + # load model on device + compiled_model = core.compile_model(model=conv_model, device_name=device.value) + +Verify Model Inference `⇑ <#top>`__ +############################################################################################################################### + Using the compiled model, we run inference on the same sample video and print the top 5 predictions again. diff --git a/docs/notebooks/211-speech-to-text-with-output.rst b/docs/notebooks/211-speech-to-text-with-output.rst index fc8fbffd21d..080d8b092c9 100644 --- a/docs/notebooks/211-speech-to-text-with-output.rst +++ b/docs/notebooks/211-speech-to-text-with-output.rst @@ -1,9 +1,11 @@ Speech to Text with OpenVINO™ ============================= +.. _top: + This tutorial demonstrates speech-to-text recognition with OpenVINO. -This tutorial uses the `quartznet +This tutorial uses the `QuartzNet 15x5 `__ model. QuartzNet performs automatic speech recognition. Its design is based on the Jasper architecture, which is a convolutional model trained @@ -11,8 +13,37 @@ with Connectionist Temporal Classification (CTC) loss. The model is available from `Open Model Zoo `__. -Imports -------- +**Table of contents**: + +- `Imports <#imports>`__ +- `Settings <#settings>`__ +- `Download and Convert Public Model <#download-and-convert-public-model>`__ + + - `Download Model <#download-model>`__ + - `Convert Model <#convert-model>`__ + +- `Audio Processing <#audio-processing>`__ + + - `Define constants <#define-constants>`__ + - `Available Audio Formats <#available-audio-formats>`__ + - `Load Audio File <#load-audio-file>`__ + - `Visualize Audio File <#visualize-audio-file>`__ + - `Change Type of Data <#change-type-of-data>`__ + - `Convert Audio to Mel Spectrum <#convert-audio-to-mel-spectrum>`__ + - `Run Conversion from Audio to Mel Format <#run-conversion-from-audio-to-mel-format>`__ + - `Visualize Mel Spectrogram <#visualize-mel-spectrogram>`__ + - `Adjust Mel scale to Input <#adjust-mel-scale-to-input>`__ + +- `Load the Model <#load-the-model>`__ + + - `Do Inference <#do-inference>`__ + - `Read Output <#read-output>`__ + - `Implementation of Decoding <#implementation-of-decoding>`__ + - `Run Decoding and Print Output <#run-decoding-and-print-output>`__ + +Imports `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -34,8 +65,9 @@ Imports from openvino.runtime import Core, serialize, Tensor from openvino.tools import mo -Settings --------- +Settings `⇑ <#top>`__ +############################################################################################################################### + In this part, all variables used in the notebook are set. @@ -48,15 +80,16 @@ In this part, all variables used in the notebook are set. precision = "FP16" model_name = "quartznet-15x5-en" -Download and Convert Public Model ---------------------------------- +Download and Convert Public Model `⇑ <#top>`__ +############################################################################################################################### -If it is your first run, models will be downloaded and converted here. -It my take a few minutes. Use ``omz_downloader`` and ``omz_converter``, -which are command-line tools from the ``openvino-dev`` package. +If it is your first run, models will be downloaded and converted here. It my take a few minutes. +Use ``omz_downloader`` and ``omz_converter``, which are command-line +tools from the ``openvino-dev`` package. + +Download Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Download Model -~~~~~~~~~~~~~~ The ``omz_downloader`` tool automatically creates a directory structure and downloads the selected model. This step is skipped if the model is @@ -74,8 +107,9 @@ Representation (OpenVINO IR). download_command = f"omz_downloader --name {model_name} --output_dir {download_folder} --precision {precision}" ! $download_command -Convert Model -~~~~~~~~~~~~~ +Convert Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + In previous step, model was downloaded in PyTorch format. Currently, PyTorch models supported in OpenVINO via ONNX exporting, @@ -176,7 +210,7 @@ Intermediate Representation format for applying optimizations. output_names=['output'], dynamic_axes={"audio_signal": {0: "batch_size", 2: "wave_len"}, "output": {0: "batch_size", 2: "wave_len"}} ) - # convert model to OpenVINO Model using OpenVINO Model Optimizer + # convert model to OpenVINO Model using model conversion API ov_model = mo.convert_model(str(onnx_model_path)) # serialize model to IR for next usage serialize(ov_model, str(converted_model_path)) @@ -191,13 +225,15 @@ Intermediate Representation format for applying optimizations. downloaded_model_path = Path("output/public/quartznet-15x5-en/models") convert_model(downloaded_model_path, path_to_converted_model) -Audio Processing ----------------- +Audio Processing `⇑ <#top>`__ +############################################################################################################################### + Now that the model is converted, load an audio file. -Define constants -~~~~~~~~~~~~~~~~ +Define constants `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + First, locate an audio file and define the alphabet used by the model. This tutorial uses the Latin alphabet beginning with a space symbol and @@ -209,17 +245,21 @@ could be any other character. audio_file_name = "edge_to_cloud.ogg" alphabet = " abcdefghijklmnopqrstuvwxyz'~" -Available Audio Formats -~~~~~~~~~~~~~~~~~~~~~~~ +Available Audio Formats `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + There are multiple supported audio formats that can be used with the model: -AIFF, AU, AVR, CAF, FLAC, HTK, SVX, MAT4, MAT5, MPC2K, OGG, PAF, PVF, -RAW, RF64, SD2, SDS, IRCAM, VOC, W64, WAV, NIST, WAVEX, WVE, XI +``AIFF``, ``AU``, ``AVR``, ``CAF``, ``FLAC``, ``HTK``, ``SVX``, +``MAT4``, ``MAT5``, ``MPC2K``, ``OGG``, ``PAF``, ``PVF``, ``RAW``, +``RF64``, ``SD2``, ``SDS``, ``IRCAM``, ``VOC``, ``W64``, ``WAV``, +``NIST``, ``WAVEX``, ``WVE``, ``XI`` + +Load Audio File `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Load Audio File -~~~~~~~~~~~~~~~ Load the file after checking a file extension. Pass ``sr`` (stands for a ``sampling rate``) as an additional parameter. The model supports files @@ -249,8 +289,9 @@ Now, you can play your audio file. -Visualise Audio File -~~~~~~~~~~~~~~~~~~~~ +Visualize Audio File `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + You can visualize how your audio file presents on a wave plot and spectrogram. @@ -286,8 +327,9 @@ spectrogram. .. image:: 211-speech-to-text-with-output_files/211-speech-to-text-with-output_21_3.png -Change Type of Data -~~~~~~~~~~~~~~~~~~~ +Change Type of Data `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The file loaded in the previous step may contain data in ``float`` type with a range of values between -1 and 1. To generate a viable input, @@ -300,8 +342,9 @@ multiply each value by the max value of ``int16`` and convert it to audio = (audio * (2**15 - 1)) audio = audio.astype(np.int16) -Convert Audio to Mel Spectrum -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Convert Audio to Mel Spectrum `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Next, convert the pre-pre-processed audio to `Mel Spectrum `__. @@ -340,8 +383,9 @@ article `__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + In this step, convert a current audio file into `Mel scale `__. @@ -350,8 +394,9 @@ scale `__. mel_basis, spec = audio_to_mel(audio=audio.flatten(), sampling_rate=sampling_rate) -Visualise Mel Spectogram -~~~~~~~~~~~~~~~~~~~~~~~~ +Visualize Mel Spectrogram `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + For more information about Mel spectrogram, refer to this `article `__. @@ -374,8 +419,9 @@ presents filter bank for converting Hz to Mels. .. image:: 211-speech-to-text-with-output_files/211-speech-to-text-with-output_29_1.png -Adjust Mel scale to Input -~~~~~~~~~~~~~~~~~~~~~~~~~ +Adjust Mel scale to Input `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Before reading the network, make sure that the input is ready. @@ -383,8 +429,9 @@ Before reading the network, make sure that the input is ready. audio = mel_to_input(mel_basis=mel_basis, spec=spec) -Load the Model --------------- +Load the Model `⇑ <#top>`__ +############################################################################################################################### + Now, you can read and load the network. @@ -392,9 +439,9 @@ Now, you can read and load the network. ie = Core() -You may run the network on multiple devices. By default, it will load -the model on CPU (you can choose manually CPU, GPU, etc.) or let -the engine choose the best available device (AUTO). +You may run the model on multiple devices. By default, it will load the +model on CPU (you can choose manually CPU, GPU etc.) or let the engine +choose the best available device (AUTO). To list all available devices that can be used, run ``print(ie.available_devices)`` command. @@ -409,9 +456,22 @@ To list all available devices that can be used, run ['CPU', 'GPU'] -To change the device used for your network, change value of -``device_name`` variable to one of the values listed by ``print()`` in -the cell above. +Select device from dropdown list + +.. code:: ipython3 + + import ipywidgets as widgets + + core = Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device .. code:: ipython3 @@ -422,10 +482,11 @@ the cell above. shape = model_input_layer.partial_shape shape[2] = -1 model.reshape({model_input_layer: shape}) - compiled_model = ie.compile_model(model=model, device_name="CPU") + compiled_model = ie.compile_model(model=model, device_name=device.value) + +Do Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Do Inference -~~~~~~~~~~~~ Everything is set up. Now, the only thing that remains is passing input to the previously loaded network and running inference. @@ -436,13 +497,14 @@ to the previously loaded network and running inference. character_probabilities = compiled_model([Tensor(audio)])[output_layer_ir] -Read Output -~~~~~~~~~~~ +Read Output `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + After inference, you need to reach out the output. The default output -format for ``quartznet 15x5`` are per-frame probabilities (after +format for ``QuartzNet 15x5`` are per-frame probabilities (after LogSoftmax) for every symbol in the alphabet, name - output, shape - -1x64x29, output data format is BxNxC, where: +1x64x29, output data format is ``BxNxC``, where: - B - batch size - N - number of audio frames @@ -466,8 +528,9 @@ The last step is getting symbols from corresponding indexes in charlist. # Run argmax to pick most possible symbols character_probabilities = np.argmax(character_probabilities, axis=1) -Implementation of Decoding -~~~~~~~~~~~~~~~~~~~~~~~~~~ +Implementation of Decoding `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + To decode previously explained output, you need the `Connectionist Temporal Classification (CTC) @@ -485,8 +548,9 @@ function. This solution will remove consecutive letters from the output. previous_letter_id = letter_index return ''.join(transcription) -Run Decoding and Print Output -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Run Decoding and Print Output `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 diff --git a/docs/notebooks/211-speech-to-text-with-output_files/index.html b/docs/notebooks/211-speech-to-text-with-output_files/index.html index a80e5c1acd3..19fc8b722bf 100644 --- a/docs/notebooks/211-speech-to-text-with-output_files/index.html +++ b/docs/notebooks/211-speech-to-text-with-output_files/index.html @@ -1,10 +1,10 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/211-speech-to-text-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/211-speech-to-text-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/211-speech-to-text-with-output_files/


../
-211-speech-to-text-with-output_21_1.png            12-Jul-2023 00:11               21971
-211-speech-to-text-with-output_21_3.png            12-Jul-2023 00:11               87067
-211-speech-to-text-with-output_29_0.png            12-Jul-2023 00:11               50625
-211-speech-to-text-with-output_29_1.png            12-Jul-2023 00:11               10083
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/211-speech-to-text-with-output_files/


../
+211-speech-to-text-with-output_21_1.png            16-Aug-2023 01:31               21971
+211-speech-to-text-with-output_21_3.png            16-Aug-2023 01:31               87067
+211-speech-to-text-with-output_29_0.png            16-Aug-2023 01:31               50625
+211-speech-to-text-with-output_29_1.png            16-Aug-2023 01:31               10083
 

diff --git a/docs/notebooks/212-pyannote-speaker-diarization-with-output.rst b/docs/notebooks/212-pyannote-speaker-diarization-with-output.rst index 52a96057d2e..ecd9c800c0e 100644 --- a/docs/notebooks/212-pyannote-speaker-diarization-with-output.rst +++ b/docs/notebooks/212-pyannote-speaker-diarization-with-output.rst @@ -1,6 +1,8 @@ Speaker diarization =================== +.. _top: + Speaker diarization is the process of partitioning an audio stream containing human speech into homogeneous segments according to the identity of each speaker. It can enhance the readability of an automatic @@ -16,7 +18,7 @@ spoke when?” With the increasing number of broadcasts, meeting recordings and voice mail collected every year, speaker diarization has received much -attention by the speech community. Seaker diarization is an essential +attention by the speech community. Speaker diarization is an essential feature for a speech recognition system to enrich the transcription with speaker labels. @@ -37,8 +39,20 @@ card `__, `repo `__ and `paper `__. -Prerequisites -------------- +**Table of contents**: + +- `Prerequisites <#prerequisites>`__ +- `Prepare pipeline <#prepare-pipeline>`__ +- `Load test audio file <#load-test-audio-file>`__ +- `Run inference pipeline <#run-inference-pipeline>`__ +- `Convert model to OpenVINO Intermediate Representation format <#convert-model-to-openvino-intermediate-representation-format>`__ +- `Select inference device <#select-inference-device>`__ +- `Replace segmentation model with OpenVINO <#replace-segmentation-model-with-openvino>`__ +- `Run speaker diarization with OpenVINO <#run-speaker-diarization-with-openvino>`__ + +Prerequisites `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -47,31 +61,35 @@ Prerequisites .. parsed-literal:: + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 ERROR: pip's dependency resolver does not currently take into account all the packages that are installed. This behaviour is the source of the following dependency conflicts. onnx 1.14.0 requires protobuf>=3.20.2, but you have protobuf 3.20.1 which is incompatible. paddlepaddle 2.5.0rc0 requires protobuf>=3.20.2; platform_system != "Windows", but you have protobuf 3.20.1 which is incompatible. ppgan 2.1.0 requires librosa==0.8.1, but you have librosa 0.9.2 which is incompatible. - ppgan 2.1.0 requires opencv-python<=4.6.0.66, but you have opencv-python 4.8.0.74 which is incompatible. + ppgan 2.1.0 requires opencv-python<=4.6.0.66, but you have opencv-python 4.8.0.76 which is incompatible. tensorflow 2.12.0 requires protobuf!=4.21.0,!=4.21.1,!=4.21.2,!=4.21.3,!=4.21.4,!=4.21.5,<5.0.0dev,>=3.20.3, but you have protobuf 3.20.1 which is incompatible. -Prepare pipeline ----------------- +Prepare pipeline `⇑ <#top>`__ +############################################################################################################################### -Traditional Speaker Diarization systems can be generalized into a five -step process: -* **Feature extraction**: transform the raw waveform into audio features like -mel spectrogram. -* **Voice activity detection**: identify the chunks in the audio where some voice -activity was observed. As we are not interested in silence and noise, we ignore -those irrelevant chunks. -* **Speaker change detection**: identify the speaker change points in the -conversation present in the audio. -* **Speech turn representation**: encode each subchunk by creating feature representations. -* **Speech turn clustering**: cluster the subchunks based on their vector -representation. Different clustering algorithms may be applied based on the -availability of cluster count (k) and the embedding process of the previous step. +Traditional Speaker Diarization systems can be generalized into a +five-step process: + +- **Feature extraction**: transform the raw waveform into audio + features like mel spectrogram. +- **Voice activity detection**: identify the chunks in the audio where + some voice activity was observed. As we are not interested in silence + and noise, we ignore those irrelevant chunks. +- **Speaker change detection**: identify the speaker change points in + the conversation present in the audio. +- **Speech turn representation**: encode each subchunk by creating + feature representations. +- **Speech turn clustering**: cluster the subchunks based on their + vector representation. Different clustering algorithms may be applied + based on the availability of cluster count (k) and the embedding + process of the previous step. The final output will be the clusters of different subchunks from the audio stream. Each cluster can be given an anonymous identifier @@ -112,7 +130,8 @@ hub `__. .. code:: python - ## login to huggingfacehub to get access to pre-trained model + + ## login to huggingfacehub to get access to pre-trained model from huggingface_hub import notebook_login, whoami try: @@ -127,8 +146,9 @@ hub `__. pipeline = Pipeline.from_pretrained("philschmid/pyannote-speaker-diarization-endpoint") -Load test audio file --------------------- +Load test audio file `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -183,8 +203,9 @@ Load test audio file .. image:: 212-pyannote-speaker-diarization-with-output_files/212-pyannote-speaker-diarization-with-output_9_1.png -Run inference pipeline ----------------------- +Run inference pipeline `⇑ <#top>`__ +############################################################################################################################### + For running inference, we should provide a path to input audio to the pipeline @@ -205,7 +226,7 @@ pipeline .. parsed-literal:: - Diarization pipeline took 15.65 s + Diarization pipeline took 15.37 s The result of running the pipeline can be represented as a diagram @@ -244,8 +265,8 @@ We can also print each time frame and corresponding speaker: start=27.8s stop=29.5s speaker_SPEAKER_02 -Convert model to OpenVINO Intermediate Representation format ------------------------------------------------------------- +Convert model to OpenVINO Intermediate Representation format. `⇑ <#top>`__ +############################################################################################################################### For best results with OpenVINO, it is recommended to convert the model to OpenVINO IR format. OpenVINO supports PyTorch via ONNX conversion. We @@ -284,8 +305,37 @@ with ``openvino.runtime.serialize``. Model successfully converted to IR and saved to pyannote-segmentation.xml -Replace segmentation model with OpenVINO ----------------------------------------- +Select inference device `⇑ <#top>`__ +############################################################################################################################### + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +Replace segmentation model with OpenVINO `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -293,7 +343,7 @@ Replace segmentation model with OpenVINO core = Core() - ov_seg_model = core.compile_model(ov_speaker_segmentation) + ov_seg_model = core.compile_model(ov_speaker_segmentation, device.value) infer_request = ov_seg_model.create_infer_request() ov_seg_out = ov_seg_model.output(0) @@ -316,8 +366,9 @@ Replace segmentation model with OpenVINO pipeline._segmentation.infer = infer_segm -Run speaker diarization with OpenVINO -------------------------------------- +Run speaker diarization with OpenVINO `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -332,7 +383,7 @@ Run speaker diarization with OpenVINO .. parsed-literal:: - Diarization pipeline took 14.98 s + Diarization pipeline took 14.69 s .. code:: ipython3 @@ -342,7 +393,7 @@ Run speaker diarization with OpenVINO -.. image:: 212-pyannote-speaker-diarization-with-output_files/212-pyannote-speaker-diarization-with-output_25_0.png +.. image:: 212-pyannote-speaker-diarization-with-output_files/212-pyannote-speaker-diarization-with-output_27_0.png diff --git a/docs/notebooks/212-pyannote-speaker-diarization-with-output_files/212-pyannote-speaker-diarization-with-output_25_0.png b/docs/notebooks/212-pyannote-speaker-diarization-with-output_files/212-pyannote-speaker-diarization-with-output_27_0.png similarity index 100% rename from docs/notebooks/212-pyannote-speaker-diarization-with-output_files/212-pyannote-speaker-diarization-with-output_25_0.png rename to docs/notebooks/212-pyannote-speaker-diarization-with-output_files/212-pyannote-speaker-diarization-with-output_27_0.png diff --git a/docs/notebooks/212-pyannote-speaker-diarization-with-output_files/index.html b/docs/notebooks/212-pyannote-speaker-diarization-with-output_files/index.html index 6d48cac5bc2..620ab266219 100644 --- a/docs/notebooks/212-pyannote-speaker-diarization-with-output_files/index.html +++ b/docs/notebooks/212-pyannote-speaker-diarization-with-output_files/index.html @@ -1,9 +1,9 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/212-pyannote-speaker-diarization-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/212-pyannote-speaker-diarization-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/212-pyannote-speaker-diarization-with-output_files/


../
-212-pyannote-speaker-diarization-with-output_14..> 12-Jul-2023 00:11                7969
-212-pyannote-speaker-diarization-with-output_25..> 12-Jul-2023 00:11                7960
-212-pyannote-speaker-diarization-with-output_9_..> 12-Jul-2023 00:11               43095
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/212-pyannote-speaker-diarization-with-output_files/


../
+212-pyannote-speaker-diarization-with-output_14..> 16-Aug-2023 01:31                7969
+212-pyannote-speaker-diarization-with-output_27..> 16-Aug-2023 01:31                7960
+212-pyannote-speaker-diarization-with-output_9_..> 16-Aug-2023 01:31               43095
 

diff --git a/docs/notebooks/213-question-answering-with-output.rst b/docs/notebooks/213-question-answering-with-output.rst index edba0969746..e3fc0ee6c8d 100644 --- a/docs/notebooks/213-question-answering-with-output.rst +++ b/docs/notebooks/213-question-answering-with-output.rst @@ -1,16 +1,41 @@ Interactive question answering with OpenVINO™ ============================================= +.. _top: + This demo shows interactive question answering with OpenVINO, using `small BERT-large-like model `__ distilled and quantized to ``INT8`` on SQuAD v1.1 training set from larger BERT-large model. The model comes from `Open Model Zoo `__. Final part -of this notebook provides live inference results from your inputs. +of this notebook provides live inference results from your inputs. + +**Table of contents**: + +- `Imports <#imports>`__ + +- `The model <#the-model>`__ + + - `Download the model <#download-the-model>`__ + - `Load the model <#load-the-model>`__ + + - `Select inference device <#select-inference-device>`__ + +- `Processing <#processing>`__ + + - `Preprocessing <#preprocessing>`__ + - `Postprocessing <#postprocessing>`__ + - `Main Processing Function <#main-processing-function>`__ + +- `Run <#run>`__ + + - `Run on local paragraphs <#run-on-local-paragraphs>`__ + - `Run on websites <#run-on-websites>`__ + +Imports `⇑ <#top>`__ +############################################################################################################################### -Imports -------- .. code:: ipython3 @@ -24,11 +49,13 @@ Imports import html_reader as reader import tokens_bert as tokens -The model ---------- +The model `⇑ <#top>`__ +############################################################################################################################### + + +Download the model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Download the model -~~~~~~~~~~~~~~~~~~ Use ``omz_downloader``, which is a command-line tool from the ``openvino-dev`` package. The ``omz_downloader`` tool automatically @@ -82,8 +109,9 @@ there is no need to use ``omz_converter``. -Load the model -~~~~~~~~~~~~~~ +Load the model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Downloaded models are located in a fixed structure, which indicates a vendor, a model name and a precision. Only a few lines of code are @@ -98,8 +126,40 @@ You can choose ``CPU`` or ``GPU`` for this model. core = Core() # Read the network and corresponding weights from a file. model = core.read_model(model_path) - # Load the model on CPU (you can use GPU as well). - compiled_model = core.compile_model(model=model, device_name="CPU") + +Select inference device `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + core = Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + compiled_model = core.compile_model(model=model, device_name=device.value) # Get input and output names of nodes. input_keys = list(compiled_model.inputs) @@ -126,8 +186,9 @@ for BERT-large-like model. -Processing ----------- +Processing `⇑ <#top>`__ +############################################################################################################################### + NLP models usually take a list of tokens as a standard input. A token is a single word converted to some integer. To provide the proper input, @@ -164,8 +225,9 @@ content from provided URLs. # Produce one big context string. return "\n".join(paragraphs) -Preprocessing -~~~~~~~~~~~~~ +Preprocessing `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The input size in this case is 384 tokens long. The main input (``input_ids``) to used BERT model consists of two parts: question @@ -247,8 +309,9 @@ documentation `__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The results from the network are raw (logits). Use the softmax function to get the probability distribution. Then, find the best answer in the @@ -340,8 +403,9 @@ answer should come with the highest score. # Return the part of the context, which is already an answer. return context[answer[1]:answer[2]], answer[0] -Main Processing Function -~~~~~~~~~~~~~~~~~~~~~~~~ +Main Processing Function `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Run question answering on a specific knowledge base (websites) and iterate through the questions. @@ -382,11 +446,13 @@ iterate through the questions. print(f"Score: {score:.2f}") print(f"Time: {end_time - start_time:.2f}s") -Run ---- +Run `⇑ <#top>`__ +############################################################################################################################### + + +Run on local paragraphs `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Run on local paragraphs -~~~~~~~~~~~~~~~~~~~~~~~ Change sources to your own to answer your questions. You can use as many sources as you want. Usually, you need to wait a few seconds for the @@ -401,9 +467,13 @@ theory `__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -You can also provide urls. Note that the context (a knowledge base) is + +You can also provide URLs. Note that the context (a knowledge base) is built from paragraphs on websites. If some information is outside the -paragraphs, the algorithm wil not be able to find it. +paragraphs, the algorithm will not be able to find it. Sample source: `OpenVINO wiki `__ Sample questions: -- What does OpenVINO mean? -- What is the license for OpenVINO? -- Where can you deploy OpenVINO code? +- What does OpenVINO mean? +- What is the license for OpenVINO? +- Where can you deploy OpenVINO code? If you want to stop the processing just put an empty string. diff --git a/docs/notebooks/214-grammar-correction-with-output.rst b/docs/notebooks/214-grammar-correction-with-output.rst index 7cea358cd1f..eaff3b6e620 100644 --- a/docs/notebooks/214-grammar-correction-with-output.rst +++ b/docs/notebooks/214-grammar-correction-with-output.rst @@ -1,6 +1,8 @@ Grammatical Error Correction with OpenVINO ========================================== +.. _top: + AI-based auto-correction products are becoming increasingly popular due to their ease of use, editing speed, and affordability. These products improve the quality of written text in emails, blogs, and chats. @@ -39,10 +41,23 @@ It consists of the following steps: - Download and convert models from a public source using the `OpenVINO integration with Hugging Face Optimum `__. -- Create an inference pipeline for grammatical error checking +- Create an inference pipeline for grammatical error checking + +**Table of contents**: + +- `How does it work? <#how-does-it-work>`__ +- `Prerequisites <#prerequisites>`__ +- `Download and Convert Models <#download-and-convert-models>`__ + + - `Select inference device <#select-inference-device>`__ + - `Grammar Checker <#grammar-checker>`__ + - `Grammar Corrector <#grammar-corrector>`__ + +- `Prepare Demo Pipeline <#prepare-demo-pipeline>`__ + +How does it work? `⇑ <#top>`__ +############################################################################################################################### -How does it work? ------------------ A Grammatical Error Correction task can be thought of as a sequence-to-sequence task where a model is trained to take a @@ -62,10 +77,10 @@ paper discovers that overall instruction finetuning is a general method that improves the performance and usability of pre-trained language models. -.. figure:: https://s3.amazonaws.com/moonup/production/uploads/1666363435475-62441d1d9fdefb55a0b7d12c.png - :alt: flan-t5 training +.. figure:: https://production-media.paperswithcode.com/methods/a04cb14e-e6b8-449e-9487-bc4262911d74.png + :alt: flan-t5_training - flan-t5 training + flan-t5_training For more details about the model, please check out `paper `__, original @@ -91,8 +106,9 @@ documentation `__ Now that we know more about FLAN-T5 and RoBERTa, let us get started. 🚀 -Prerequisites -------------- +Prerequisites `⇑ <#top>`__ +############################################################################################################################### + First, we need to install the `Hugging Face Optimum `__ library @@ -106,15 +122,25 @@ documentation `__. !pip install -q "git+https://github.com/huggingface/optimum-intel.git" onnx onnxruntime -Download and Convert Models ---------------------------- + +.. parsed-literal:: + + + [notice] A new release of pip is available: 23.1.2 -> 23.2 + [notice] To update, run: pip install --upgrade pip + + +Download and Convert Models `⇑ <#top>`__ +############################################################################################################################### + Optimum Intel can be used to load optimized models from the `Hugging Face Hub `__ and create pipelines to run an inference with OpenVINO Runtime using Hugging Face APIs. The Optimum Inference models are API compatible with Hugging Face Transformers models. This means we just need to replace -AutoModelForXxx class with the corresponding OVModelForXxx class. +``AutoModelForXxx`` class with the corresponding ``OVModelForXxx`` +class. Below is an example of the RoBERTa text classification model @@ -143,9 +169,10 @@ Tokenizer class and pipelines API are compatible with Optimum models. .. parsed-literal:: - 2023-02-22 08:52:28.563283: I tensorflow/core/util/util.cc:169] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. - /home/ea/work/my_ov/openvino/tmp_notebooks_env/lib/python3.8/site-packages/openvino/offline_transformations/__init__.py:10: FutureWarning: The module is private and following namespace `offline_transformations` will be removed in the future, use `openvino.runtime.passes` instead! - warnings.warn( + 2023-07-17 14:43:08.812267: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-07-17 14:43:08.850959: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. + 2023-07-17 14:43:09.468643: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT .. parsed-literal:: @@ -156,10 +183,43 @@ Tokenizer class and pipelines API are compatible with Optimum models. .. parsed-literal:: No CUDA runtime is found, using CUDA_HOME='/usr/local/cuda' + comet_ml is installed but `COMET_API_KEY` is not set. -Grammar Checker -~~~~~~~~~~~~~~~ +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + from openvino.runtime import Core + + core = Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=2, options=('CPU', 'GPU', 'AUTO'), value='AUTO') + + + +Grammar Checker `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -168,11 +228,18 @@ Grammar Checker grammar_checker_tokenizer = AutoTokenizer.from_pretrained(grammar_checker_model_id) if grammar_checker_dir.exists(): - grammar_checker_model = OVModelForSequenceClassification.from_pretrained(grammar_checker_dir) + grammar_checker_model = OVModelForSequenceClassification.from_pretrained(grammar_checker_dir, device=device.value) else: - grammar_checker_model = OVModelForSequenceClassification.from_pretrained(grammar_checker_model_id, from_transformers=True) + grammar_checker_model = OVModelForSequenceClassification.from_pretrained(grammar_checker_model_id, export=True, device=device.value) grammar_checker_model.save_pretrained(grammar_checker_dir) + +.. parsed-literal:: + + Compiling the model... + Set CACHE_DIR to roberta-base-cola/model_cache + + Let us check model work, using inference pipeline for ``text-classification`` task. You can find more information about usage Hugging Face inference pipelines in this @@ -188,6 +255,12 @@ Hugging Face inference pipelines in this print(f'predicted score: {result["score"] :.2}') +.. parsed-literal:: + + Xformers is not installed correctly. If you want to use memory_efficient_attention to accelerate training use the following command to install Xformers + pip install xformers. + + .. parsed-literal:: input text: They are moved by salar energy @@ -197,8 +270,9 @@ Hugging Face inference pipelines in this Great! Looks like the model can detect errors in the sample. -Grammar Corrector -~~~~~~~~~~~~~~~~~ +Grammar Corrector `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The steps for loading the Grammar Corrector model are very similar, except for the model class that is used. Because FLAN-T5 is a @@ -213,16 +287,34 @@ to run it. grammar_corrector_tokenizer = AutoTokenizer.from_pretrained(grammar_corrector_model_id) if grammar_corrector_dir.exists(): - grammar_corrector_model = OVModelForSeq2SeqLM.from_pretrained(grammar_corrector_dir) + grammar_corrector_model = OVModelForSeq2SeqLM.from_pretrained(grammar_corrector_dir, device=device.value) else: - grammar_corrector_model = OVModelForSeq2SeqLM.from_pretrained(grammar_corrector_model_id, from_transformers=True) + grammar_corrector_model = OVModelForSeq2SeqLM.from_pretrained(grammar_corrector_model_id, export=True, device=device.value) grammar_corrector_model.save_pretrained(grammar_corrector_dir) .. parsed-literal:: + The argument `from_transformers` is deprecated, and will be removed in optimum 2.0. Use `export` instead + Framework not specified. Using pt to export to ONNX. + Using framework PyTorch: 1.13.1+cpu + Overriding 1 configuration item(s) + - use_cache -> False + Using framework PyTorch: 1.13.1+cpu + Overriding 1 configuration item(s) + - use_cache -> True + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/transformers/modeling_utils.py:850: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + if causal_mask.shape[1] < attention_mask.shape[1]: + Using framework PyTorch: 1.13.1+cpu + Overriding 1 configuration item(s) + - use_cache -> True + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/transformers/models/t5/modeling_t5.py:507: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + elif past_key_value.shape[2] != key_value_states.shape[1]: In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + Compiling the encoder... + Compiling the decoder... + Compiling the decoder... .. code:: ipython3 @@ -244,8 +336,9 @@ to run it. Nice! The result looks pretty good! -Prepare Demo Pipeline ---------------------- +Prepare Demo Pipeline `⇑ <#top>`__ +############################################################################################################################### + Now let us put everything together and create the pipeline for grammar correction. The pipeline accepts input text, verifies its correctness, @@ -366,6 +459,14 @@ execute the following cells. corrected_text = correct_text(text_widget.value, grammar_checker_pipe, grammar_corrector_pipe) +.. parsed-literal:: + + huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks... + To disable this warning, you can either: + - Avoid using `tokenizers` before the fork if possible + - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false) + + .. parsed-literal:: diff --git a/docs/notebooks/215-image-inpainting-with-output.rst b/docs/notebooks/215-image-inpainting-with-output.rst index 38047dcff54..85f762359ab 100644 --- a/docs/notebooks/215-image-inpainting-with-output.rst +++ b/docs/notebooks/215-image-inpainting-with-output.rst @@ -1,6 +1,8 @@ Image In-painting with OpenVINO™ -------------------------------- +.. _top: + This notebook demonstrates how to use an image in-painting model with OpenVINO, using `GMCNN model `__ from `Open Model @@ -9,6 +11,19 @@ given a tampered image, is able to create something very similar to the original image. The Following pipeline will be used in this notebook. |pipeline| +**Table of contents**: + +- `Download the Model <#download-the-model>`__ +- `Convert Tensorflow model to OpenVINO IR format <#convert-tensorflow-model-to-openvino-ir-format>`__ +- `Load the model <#load-the-model>`__ +- `Determine the input shapes of the model <#determine-the-input-shapes-of-the-model>`__ +- `Create a square mask <#create-a-square-mask>`__ +- `Load and Resize the Image <#load-and-resize-the-image>`__ +- `Generating the Masked Image <#generating-the-masked-image>`__ +- `Preprocessing <#preprocessing>`__ +- `Inference <#inference>`__ +- `Save the Restored Image <#save-the-restored-image>`__ + .. |pipeline| image:: https://user-images.githubusercontent.com/4547501/165792473-ba784c0d-0a37-409f-a5f6-bb1849c1d140.png .. code:: ipython3 @@ -26,13 +41,13 @@ original image. The Following pipeline will be used in this notebook. sys.path.append("../utils") import notebook_utils as utils -Download the Model -~~~~~~~~~~~~~~~~~~ +Download the Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Download ``gmcnn-places2-tf``\ model (this step will be skipped if the -model is already downloaded) and then unzip it. Downloaded model stored -in TensorFlow frozen graph format. The steps how this frozen graph can -be obtained from original model checkpoint can be found in this +Download ``gmcnn-places2-tf``\ model (this step will be skipped if the model is already downloaded) and then +unzip it. Downloaded model stored in TensorFlow frozen graph format. The +steps how this frozen graph can be obtained from original model +checkpoint can be found in this `instruction `__ .. code:: ipython3 @@ -58,14 +73,14 @@ be obtained from original model checkpoint can be found in this Already downloaded -Convert Tensorflow model to OpenVINO IR format -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Convert Tensorflow model to OpenVINO IR format `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The pre-trained model is in TensorFlow format. To use it with OpenVINO, -convert it to OpenVINO IR format. To do this Model Optimizer is used. -For more information about Model Optimizer, see the `Model Optimizer -Developer -Guide `__. +convert it to OpenVINO IR format with model conversion API. For more +information about model conversion, see this +`page `__. This step is also skipped if the model is already converted. .. code:: ipython3 @@ -73,7 +88,7 @@ This step is also skipped if the model is already converted. model_dir = Path(base_model_dir, 'public', 'ir') ir_path = Path(f"{model_dir}/frozen_model.xml") - # Run Model Optimizer to convert model to OpenVINO IR FP32 format, if the IR file does not exist. + # Run model conversion API to convert model to OpenVINO IR FP32 format, if the IR file does not exist. if not ir_path.exists(): ov_model = mo.convert_model(model_path, input_shape=[[1,512,680,3],[1,512,680,1]]) serialize(ov_model, str(ir_path)) @@ -86,8 +101,9 @@ This step is also skipped if the model is already converted. model/public/ir/frozen_model.xml already exists. -Load the model -~~~~~~~~~~~~~~ +Load the model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Now, load the OpenVINO IR model and perform as follows: @@ -105,14 +121,40 @@ Only a few lines of code are required to run the model: # Read the model.xml and weights file model = core.read_model(model=ir_path) - # Load the model on to the CPU - compiled_model = core.compile_model(model=model, device_name="CPU") + +.. code:: ipython3 + + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=2, options=('CPU', 'GPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + # Load the model on to the device + compiled_model = core.compile_model(model=model, device_name=device.value) # Store the input and output nodes input_layer = compiled_model.input(0) output_layer = compiled_model.output(0) -Determine the input shapes of the model -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Determine the input shapes of the model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Note that both input shapes are the same. However, the second input has 1 channel (monotone). @@ -121,8 +163,9 @@ Note that both input shapes are the same. However, the second input has N, H, W, C = input_layer.shape -Create a square mask -~~~~~~~~~~~~~~~~~~~~ +Create a square mask `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Next, create a single channeled mask that will be laid on top of the original image. @@ -161,11 +204,12 @@ original image. -.. image:: 215-image-inpainting-with-output_files/215-image-inpainting-with-output_12_0.png +.. image:: 215-image-inpainting-with-output_files/215-image-inpainting-with-output_14_0.png -Load and Resize the Image -~~~~~~~~~~~~~~~~~~~~~~~~~ +Load and Resize the Image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + This image will be altered by using the mask. You can process any image you like. Just change the URL below. @@ -190,11 +234,12 @@ you like. Just change the URL below. -.. image:: 215-image-inpainting-with-output_files/215-image-inpainting-with-output_14_0.png +.. image:: 215-image-inpainting-with-output_files/215-image-inpainting-with-output_16_0.png -Generating the Masked Image -~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Generating the Masked Image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + This multiplication of the image and the mask gives the result of the masked image layered on top of the original image. The ``masked_image`` @@ -209,11 +254,12 @@ will be the first input to the GMCNN model. -.. image:: 215-image-inpainting-with-output_files/215-image-inpainting-with-output_16_0.png +.. image:: 215-image-inpainting-with-output_files/215-image-inpainting-with-output_18_0.png -Preprocessing -~~~~~~~~~~~~~ +Preprocessing `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The model expects the input dimensions to be ``NHWC``. @@ -225,8 +271,9 @@ The model expects the input dimensions to be ``NHWC``. masked_image = masked_image[None, ...] mask = mask[None, ...] -Inference -~~~~~~~~~ +Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Do inference with the given masked image and the mask. Then, show the restored image. @@ -240,11 +287,12 @@ restored image. -.. image:: 215-image-inpainting-with-output_files/215-image-inpainting-with-output_20_0.png +.. image:: 215-image-inpainting-with-output_files/215-image-inpainting-with-output_22_0.png -Save the Restored Image -~~~~~~~~~~~~~~~~~~~~~~~ +Save the Restored Image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Save the restored image to the data directory to download it. diff --git a/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_12_0.png b/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_12_0.png deleted file mode 100644 index 4b9e9d29605..00000000000 --- a/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_12_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:a7a3868211d34f6eb06edce30cb80a949216610c58cb5399e9c17691a8453a5a -size 16222 diff --git a/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_14_0.png b/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_14_0.png index 0ffd97041ff..5e0d90c8fca 100644 --- a/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_14_0.png +++ b/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_14_0.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:9b69904bb953a31d2c89974a5ae14753aa290e0e197a76d8c160442ec995846b -size 544222 +oid sha256:a46cd32b28865daa8ce534fd66b37766ddbb2c25806469b3681d33df0b671f18 +size 16155 diff --git a/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_16_0.png b/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_16_0.png index b9939689142..0ffd97041ff 100644 --- a/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_16_0.png +++ b/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_16_0.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:cda42e8dbb2992ec7644d01f1ec1d8614e4747d85102ca3e294d7edba21ccd54 -size 508083 +oid sha256:9b69904bb953a31d2c89974a5ae14753aa290e0e197a76d8c160442ec995846b +size 544222 diff --git a/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_18_0.png b/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_18_0.png new file mode 100644 index 00000000000..c7242ebac0f --- /dev/null +++ b/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_18_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:81b3eed7308c70674b16172ecf146179ee5e60b431deb8d8629d6766c7fb7e50 +size 493354 diff --git a/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_20_0.png b/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_20_0.png deleted file mode 100644 index 31b4e958d45..00000000000 --- a/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_20_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:4eaeede639723ce8a680a57a25aef08db871d8c572497b8a78665f68ecca5648 -size 592799 diff --git a/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_22_0.png b/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_22_0.png new file mode 100644 index 00000000000..d56b6882f78 --- /dev/null +++ b/docs/notebooks/215-image-inpainting-with-output_files/215-image-inpainting-with-output_22_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:6177d3e806c58c63147586c10aaa45312883fcb294b2f5ac64ace11c64d29673 +size 586544 diff --git a/docs/notebooks/215-image-inpainting-with-output_files/index.html b/docs/notebooks/215-image-inpainting-with-output_files/index.html index 9722cd89d6b..1afb401da23 100644 --- a/docs/notebooks/215-image-inpainting-with-output_files/index.html +++ b/docs/notebooks/215-image-inpainting-with-output_files/index.html @@ -1,10 +1,10 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/215-image-inpainting-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/215-image-inpainting-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/215-image-inpainting-with-output_files/


../
-215-image-inpainting-with-output_12_0.png          12-Jul-2023 00:11               16222
-215-image-inpainting-with-output_14_0.png          12-Jul-2023 00:11              544222
-215-image-inpainting-with-output_16_0.png          12-Jul-2023 00:11              508083
-215-image-inpainting-with-output_20_0.png          12-Jul-2023 00:11              592799
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/215-image-inpainting-with-output_files/


../
+215-image-inpainting-with-output_14_0.png          16-Aug-2023 01:31               16155
+215-image-inpainting-with-output_16_0.png          16-Aug-2023 01:31              544222
+215-image-inpainting-with-output_18_0.png          16-Aug-2023 01:31              493354
+215-image-inpainting-with-output_22_0.png          16-Aug-2023 01:31              586544
 

diff --git a/docs/notebooks/216-attention-center-with-output.rst b/docs/notebooks/216-attention-center-with-output.rst index d79923fb459..07e5c69eedb 100644 --- a/docs/notebooks/216-attention-center-with-output.rst +++ b/docs/notebooks/216-attention-center-with-output.rst @@ -1,11 +1,13 @@ The attention center model with OpenVINO™ ========================================= +.. _top: + This notebook demonstrates how to use the `attention center model `__ with OpenVINO. This model is in the `TensorFlow Lite format `__, which is supported in -OpenVINO now by TFlite frontend. +OpenVINO now by TFLite frontend. Eye tracking is commonly used in visual neuroscience and cognitive science to answer related questions such as visual attention and @@ -16,7 +18,7 @@ output. This 2D point is the predicted center of human attention on the image i.e. the most salient part of images, on which people pay attention fist to. This allows find the most visually salient regions and handle it as early as possible. For example, it could be used for -the latest generatipon image format(such as `JPEG +the latest generation image format (such as `JPEG XL `__), which supports encoding the parts that you pay attention to fist. It can help to improve user experience, image will appear to load faster. @@ -37,22 +39,33 @@ attention center. And then an operator (the Einstein summation operator in our case) can be applied to compute the (gravity) center from the weighting map. An L2 norm between the predicted attention center and the ground-truth attention center can be computed as the training loss. -Source: `google AI -blogpost `__. +Source: `Google AI blog +post `__. -.. image:: https://camo.githubusercontent.com/6fabb912edba4b321f2ffff55235e68f8d8ccb8d90373126788ddad25fe79708/68747470733a2f2f626c6f676765722e676f6f676c6575736572636f6e74656e742e636f6d2f696d672f622f523239765a32786c2f4156765873456a784c43444a487a4a4e6a425f766f6e2d76466c7138544a4a4641343161423835542d5145335a4e7857386b7368416633484f457949454a34756767586a624a6d5a6873646a376a3669366d76766d5874796178584a506d334a48754b494c4e5254506658394b7649436246425244384b4e7544566d4c41427a597568516369334254324271562d774d35344978616f415631594442626e704a433932555a6645424776616b4c757369714e44324161507057507232674a56312f73313630302f696d616765342e706e67 +.. figure:: https://blogger.googleusercontent.com/img/b/R29vZ2xl/AVvXsEjxLCDJHzJNjB_von-vFlq8TJJFA41aB85T-QE3ZNxW8kshAf3HOEyIEJ4uggXjbJmZhsdj7j6i6mvvmXtyaxXJPm3JHuKILNRTPfX9KvICbFBRD8KNuDVmLABzYuhQci3BT2BqV-wM54IxaoAV1YDBbnpJC92UZfEBGvakLusiqND2AaPpWPr2gJV1/s1600/image4.png + :alt: drawing + + drawing The attention center model has been trained with images from the `COCO dataset `__ annotated with saliency from -the `salicon dataset `__. +the `SALICON dataset `__. -The tutorial consists of the following steps: -* Downloading the model -* Loading the model and make inference with OpenVINO API -* Running Live Attention Center Detection +**Table of contents**: + +- `Imports <#imports>`__ +- `Download the attention-center model <#download-the-attention-center-model>`__ + + - `Convert Tensorflow Lite model to OpenVINO IR format <#convert-tensorflow-lite-model-to-openvino-ir-format>`__ + +- `Select inference device <#select-inference-device>`__ +- `Prepare image to use with attention-center model <#prepare-image-to-use-with-attention-center-model>`__ +- `Load input image <#load-input-image>`__ +- `Get result with OpenVINO IR model <#get-result-with-openvino-ir-model>`__ + +Imports `⇑ <#top>`__ +############################################################################################################################### -Imports -------- .. code:: ipython3 @@ -69,14 +82,15 @@ Imports .. parsed-literal:: - 2023-07-11 23:09:56.206795: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. - 2023-07-11 23:09:56.240689: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + 2023-08-15 23:14:52.395540: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-08-15 23:14:52.429075: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. - 2023-07-11 23:09:56.780351: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + 2023-08-15 23:14:52.969814: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT -Download the attention-center model ------------------------------------ +Download the attention-center model `⇑ <#top>`__ +############################################################################################################################### + Download the model as part of `attention-center repo `__. The repo @@ -95,12 +109,13 @@ include model in folder ``./model``. remote: Counting objects: 100% (168/168), done. remote: Compressing objects: 100% (132/132), done. remote: Total 168 (delta 73), reused 114 (delta 28), pack-reused 0 - Receiving objects: 100% (168/168), 26.22 MiB | 4.23 MiB/s, done. + Receiving objects: 100% (168/168), 26.22 MiB | 4.18 MiB/s, done. Resolving deltas: 100% (73/73), done. -Convert Tensorflow Lite model to OpenVINO IR format -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Convert Tensorflow Lite model to OpenVINO IR format `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The attention-center model is pre-trained model in TensorFlow Lite format. In this Notebook the model will be converted to OpenVINO IR @@ -109,7 +124,7 @@ already been converted. For more information about Model Optimizer, please, see the `Model Optimizer Developer Guide `__. -Also TFLite models format is supported in OpenVINO by TFlite frontend, +Also TFLite models format is supported in OpenVINO by TFLite frontend, so the model can be passed directly to ``core.read_model()``. You can find example in `002-openvino-api `__. @@ -129,9 +144,6 @@ find example in else: print("Read IR model from {}".format(ir_model_path)) model = core.read_model(ir_model_path) - - device = "CPU" - compiled_model = core.compile_model(model=model, device_name=device) .. parsed-literal:: @@ -139,8 +151,41 @@ find example in IR model saved to model/ir_center_model.xml -Prepare image to use with attention-center model ------------------------------------------------- +Select inference device `⇑ <#top>`__ +############################################################################################################################### + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + compiled_model = core.compile_model(model=model, device_name=device.value) + +Prepare image to use with attention-center model `⇑ <#top>`__ +############################################################################################################################### + The attention-center model takes an RGB image with shape (480, 640) as input. @@ -190,8 +235,9 @@ input. plt.imshow(cv2.cvtColor(image_to_print, cv2.COLOR_BGR2RGB)) -Load input image ----------------- +Load input image `⇑ <#top>`__ +############################################################################################################################### + Upload input image using file loading button @@ -229,16 +275,17 @@ Upload input image using file loading button .. parsed-literal:: - 2023-07-11 23:10:08.360269: W tensorflow/core/common_runtime/gpu/gpu_device.cc:1956] Cannot dlopen some GPU libraries. Please make sure the missing libraries mentioned above are installed properly if you would like to use GPU. Follow the guide at https://www.tensorflow.org/install/gpu for how to download and setup the required libraries for your platform. + 2023-08-15 23:15:04.645356: W tensorflow/core/common_runtime/gpu/gpu_device.cc:1956] Cannot dlopen some GPU libraries. Please make sure the missing libraries mentioned above are installed properly if you would like to use GPU. Follow the guide at https://www.tensorflow.org/install/gpu for how to download and setup the required libraries for your platform. Skipping registering GPU devices... -.. image:: 216-attention-center-with-output_files/216-attention-center-with-output_11_1.png +.. image:: 216-attention-center-with-output_files/216-attention-center-with-output_14_1.png -Get result with OpenVINO IR model ---------------------------------- +Get result with OpenVINO IR model `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -258,5 +305,5 @@ Get result with OpenVINO IR model -.. image:: 216-attention-center-with-output_files/216-attention-center-with-output_13_1.png +.. image:: 216-attention-center-with-output_files/216-attention-center-with-output_16_1.png diff --git a/docs/notebooks/216-attention-center-with-output_files/216-attention-center-with-output_11_1.png b/docs/notebooks/216-attention-center-with-output_files/216-attention-center-with-output_14_1.png similarity index 100% rename from docs/notebooks/216-attention-center-with-output_files/216-attention-center-with-output_11_1.png rename to docs/notebooks/216-attention-center-with-output_files/216-attention-center-with-output_14_1.png diff --git a/docs/notebooks/216-attention-center-with-output_files/216-attention-center-with-output_13_1.png b/docs/notebooks/216-attention-center-with-output_files/216-attention-center-with-output_16_1.png similarity index 100% rename from docs/notebooks/216-attention-center-with-output_files/216-attention-center-with-output_13_1.png rename to docs/notebooks/216-attention-center-with-output_files/216-attention-center-with-output_16_1.png diff --git a/docs/notebooks/216-attention-center-with-output_files/index.html b/docs/notebooks/216-attention-center-with-output_files/index.html index 47eb139a18d..96d527c4c5f 100644 --- a/docs/notebooks/216-attention-center-with-output_files/index.html +++ b/docs/notebooks/216-attention-center-with-output_files/index.html @@ -1,8 +1,8 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/216-attention-center-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/216-attention-center-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/216-attention-center-with-output_files/


../
-216-attention-center-with-output_11_1.png          12-Jul-2023 00:11              387941
-216-attention-center-with-output_13_1.png          12-Jul-2023 00:11              387905
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/216-attention-center-with-output_files/


../
+216-attention-center-with-output_14_1.png          16-Aug-2023 01:31              387941
+216-attention-center-with-output_16_1.png          16-Aug-2023 01:31              387905
 

diff --git a/docs/notebooks/217-vision-deblur-with-output.rst b/docs/notebooks/217-vision-deblur-with-output.rst index e68f0d35492..3686de8db5f 100644 --- a/docs/notebooks/217-vision-deblur-with-output.rst +++ b/docs/notebooks/217-vision-deblur-with-output.rst @@ -1,6 +1,28 @@ Deblur Photos with DeblurGAN-v2 and OpenVINO™ ============================================= +.. _top: + +**Table of contents**: + +- `What is deblurring? <#what-is-deblurring>`__ +- `Preparations <#preparations>`__ + + - `Imports <#imports>`__ + - `Settings <#settings>`__ + - `Select inference device <#select-inference-device>`__ + - `Download DeblurGAN-v2 Model <#download-deblurgan-v2-model>`__ + - `Prepare model <#prepare-model>`__ + - `Convert DeblurGAN-v2 Model to OpenVINO IR format <#convert-deblurgan-v2-model-to-openvino-ir-format>`__ + - `Load the Model <#load-the-model>`__ + +- `Deblur Image <#deblur-image>`__ + + - `Load, resize and reshape input image <#load-resize-and-reshape-input-image>`__ + - `Do Inference on the Input Image <#do-inference-on-the-input-image>`__ + - `Display results <#display-results>`__ + - `Save the deblurred image <#save-the-deblurred-image>`__ + This tutorial demonstrates Single Image Motion Deblurring with DeblurGAN-v2 in OpenVINO, by first converting the `VITA-Group/DeblurGANv2 `__ @@ -8,8 +30,9 @@ model to OpenVINO Intermediate Representation (OpenVINO IR) format. For more information about the model, see the `documentation `__. -What is deblurring? -~~~~~~~~~~~~~~~~~~~ +What is deblurring? `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Deblurring is the task of removing motion blurs that usually occur in photos shot with hand-held cameras when there are moving objects in the @@ -18,17 +41,19 @@ the image, but also complicate computer vision analyses. For more information, refer to the following research paper: -Kupyn, O., Martyniuk, T., Wu, J., & Wang, Z. (2019). `Deblurgan-v2: +Kupyn, O., Martyniuk, T., Wu, J., & Wang, Z. (2019). `DeblurGAN-v2: Deblurring (orders-of-magnitude) faster and better. `__ In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 8878-8887). -Preparations ------------- +Preparations `⇑ <#top>`__ +############################################################################################################################### + + +Imports `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Imports -~~~~~~~ .. code:: ipython3 @@ -44,14 +69,12 @@ Imports sys.path.append("../utils") from notebook_utils import load_image -Settings -~~~~~~~~ +Settings `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 - # A device to use for inference. For example, "CPU", or "GPU". - DEVICE = "CPU" - # A directory where the model will be downloaded. model_dir = Path("model") model_dir.mkdir(exist_ok=True) @@ -63,8 +86,39 @@ Settings precision = "FP16" -Download DeblurGAN-v2 Model -~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + core = Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +Download DeblurGAN-v2 Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Model defined in `VITA-Group/DeblurGANv2 `__ @@ -119,8 +173,9 @@ Downloading deblurgan-v2… -Prepare model -~~~~~~~~~~~~~ +Prepare model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + DeblurGAN-v2 is PyTorch model for converting it to OpenVINO Intermediate Representation format, we should first instantiate model class and load @@ -151,19 +206,20 @@ checkpoint weights. out = (out + 1) / 2 return out -Convert DeblurGAN-v2 Model to OpenVINO IR format -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Convert DeblurGAN-v2 Model to OpenVINO IR format `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + For best results with OpenVINO, it is recommended to convert the model -to OpenVINO IR format. We will use Model Optimizer Python API -functionality to convert the PyTorch model. The ``mo.convert_model`` -Python function returns an OpenVINO model ready to load on device and -start making predictions. We can save it on disk for next usage with -``openvino.runtime.serialize``. For more information about Model -Optimizer Python API, see the `Model Optimizer Developer -Guide `__. +to OpenVINO IR format. To convert the PyTorch model, we will use model +conversion Python API. The ``mo.convert_model`` Python function returns +an OpenVINO model ready to load on a device and start making +predictions. We can save it on a disk for next usage with +``openvino.runtime.serialize``. For more information about model +conversion Python API, see this +`page `__. -Model Conversion may take a while. +Model conversion may take a while. .. code:: ipython3 @@ -177,8 +233,9 @@ Model Conversion may take a while. ov_model = mo.convert_model(deblur_gan_model, input_shape=[[1,3,736,1312]], compress_to_fp16=(precision == "FP16")) serialize(ov_model, model_xml_path) -Load the Model --------------- +Load the Model `⇑ <#top>`__ +############################################################################################################################### + Load and compile the DeblurGAN-v2 model in the OpenVINO Runtime with ``ie.read_model`` and compile it for the specified device with @@ -189,7 +246,7 @@ shape for the model. ie = Core() model = ie.read_model(model=model_xml_path) - compiled_model = ie.compile_model(model=model, device_name=DEVICE) + compiled_model = ie.compile_model(model=model, device_name=device.value) .. code:: ipython3 @@ -222,11 +279,13 @@ shape for the model. -Deblur Image ------------- +Deblur Image `⇑ <#top>`__ +############################################################################################################################### + + +Load, resize and reshape input image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Load, resize and reshape input image -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The input image is read by using the default ``load_image`` function from ``notebooks.utils``. Then, resized to meet the network expected @@ -268,11 +327,12 @@ height, and ``W`` is the width. -.. image:: 217-vision-deblur-with-output_files/217-vision-deblur-with-output_22_0.png +.. image:: 217-vision-deblur-with-output_files/217-vision-deblur-with-output_24_0.png -Do Inference on the Input Image -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Do Inference on the Input Image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Do the inference, convert the result to an image shape and resize it to the original image size. @@ -296,11 +356,12 @@ the original image size. -.. image:: 217-vision-deblur-with-output_files/217-vision-deblur-with-output_25_0.png +.. image:: 217-vision-deblur-with-output_files/217-vision-deblur-with-output_27_0.png -Display results -~~~~~~~~~~~~~~~ +Display results `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -317,11 +378,12 @@ Display results -.. image:: 217-vision-deblur-with-output_files/217-vision-deblur-with-output_27_0.png +.. image:: 217-vision-deblur-with-output_files/217-vision-deblur-with-output_29_0.png -Save the deblurred image -~~~~~~~~~~~~~~~~~~~~~~~~ +Save the deblurred image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Save the output image of the DeblurGAN-v2 model in the current directory. diff --git a/docs/notebooks/217-vision-deblur-with-output_files/217-vision-deblur-with-output_22_0.png b/docs/notebooks/217-vision-deblur-with-output_files/217-vision-deblur-with-output_24_0.png similarity index 100% rename from docs/notebooks/217-vision-deblur-with-output_files/217-vision-deblur-with-output_22_0.png rename to docs/notebooks/217-vision-deblur-with-output_files/217-vision-deblur-with-output_24_0.png diff --git a/docs/notebooks/217-vision-deblur-with-output_files/217-vision-deblur-with-output_25_0.png b/docs/notebooks/217-vision-deblur-with-output_files/217-vision-deblur-with-output_25_0.png deleted file mode 100644 index d6f023b5f3d..00000000000 --- a/docs/notebooks/217-vision-deblur-with-output_files/217-vision-deblur-with-output_25_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:d4ae5cf320917ad4a6326089fe287a5b20986173a8b8909cb484076aabeef784 -size 223269 diff --git a/docs/notebooks/217-vision-deblur-with-output_files/217-vision-deblur-with-output_27_0.png b/docs/notebooks/217-vision-deblur-with-output_files/217-vision-deblur-with-output_27_0.png index eaea9531ac8..d6f023b5f3d 100644 --- a/docs/notebooks/217-vision-deblur-with-output_files/217-vision-deblur-with-output_27_0.png +++ b/docs/notebooks/217-vision-deblur-with-output_files/217-vision-deblur-with-output_27_0.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:8095d3c856ae1652d7a57df4a1c7e154ddb17aa2b1b53da84c40717ab4ce2da2 -size 768422 +oid sha256:d4ae5cf320917ad4a6326089fe287a5b20986173a8b8909cb484076aabeef784 +size 223269 diff --git a/docs/notebooks/217-vision-deblur-with-output_files/217-vision-deblur-with-output_29_0.png b/docs/notebooks/217-vision-deblur-with-output_files/217-vision-deblur-with-output_29_0.png new file mode 100644 index 00000000000..eaea9531ac8 --- /dev/null +++ b/docs/notebooks/217-vision-deblur-with-output_files/217-vision-deblur-with-output_29_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:8095d3c856ae1652d7a57df4a1c7e154ddb17aa2b1b53da84c40717ab4ce2da2 +size 768422 diff --git a/docs/notebooks/217-vision-deblur-with-output_files/index.html b/docs/notebooks/217-vision-deblur-with-output_files/index.html index 4fa8a8eed00..5eb34511562 100644 --- a/docs/notebooks/217-vision-deblur-with-output_files/index.html +++ b/docs/notebooks/217-vision-deblur-with-output_files/index.html @@ -1,9 +1,9 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/217-vision-deblur-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/217-vision-deblur-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/217-vision-deblur-with-output_files/


../
-217-vision-deblur-with-output_22_0.png             12-Jul-2023 00:11              220275
-217-vision-deblur-with-output_25_0.png             12-Jul-2023 00:11              223269
-217-vision-deblur-with-output_27_0.png             12-Jul-2023 00:11              768422
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/217-vision-deblur-with-output_files/


../
+217-vision-deblur-with-output_24_0.png             16-Aug-2023 01:31              220275
+217-vision-deblur-with-output_27_0.png             16-Aug-2023 01:31              223269
+217-vision-deblur-with-output_29_0.png             16-Aug-2023 01:31              768422
 

diff --git a/docs/notebooks/218-vehicle-detection-and-recognition-with-output.rst b/docs/notebooks/218-vehicle-detection-and-recognition-with-output.rst index 136252314ac..e8f4c4aafb6 100644 --- a/docs/notebooks/218-vehicle-detection-and-recognition-with-output.rst +++ b/docs/notebooks/218-vehicle-detection-and-recognition-with-output.rst @@ -1,6 +1,8 @@ Vehicle Detection And Recognition with OpenVINO™ ================================================ +.. _top: + This tutorial demonstrates how to use two pre-trained models from `Open Model Zoo `__: `vehicle-detection-0200 `__ @@ -17,10 +19,28 @@ As a result, you can get: result +**Table of contents**: + +- `Imports <#imports>`__ +- `Download Models <#download-models>`__ +- `Load Models <#load-models>`__ + + - `Get attributes from model <#get-attributes-from-model>`__ + - `Helper function <#helper-function>`__ + - `Read and display a test image <#read-and-display-a-test-image>`__ + +- `Use the Detection Model to Detect Vehicles <#use-the-detection-model-to-detect-vehicles>`__ + + - `Detection Processing <#detection-processing>`__ + - `Recognize vehicle attributes <#recognize-vehicle-attributes>`__ + - `Recognition processing <#recognition-processing>`__ + - `Combine two models <#combine-two-models>`__ + .. |flowchart| image:: https://user-images.githubusercontent.com/47499836/157867076-9e997781-f9ef-45f6-9a51-b515bbf41048.png -Imports -------- +Imports `⇑ <#top>`__ +############################################################################################################################### + Import the required modules. @@ -39,8 +59,9 @@ Import the required modules. sys.path.append("../utils") import notebook_utils as utils -Download Models ---------------- +Download Models `⇑ <#top>`__ +############################################################################################################################### + Use ``omz_downloader`` - a command-line tool from the ``openvino-dev`` package. The ``omz_downloader`` tool automatically creates a directory @@ -114,8 +135,9 @@ Representation (OpenVINO IR). -Load Models ------------ +Load Models `⇑ <#top>`__ +############################################################################################################################### + This tutorial requires a detection model and a recognition model. After downloading the models, initialize OpenVINO Runtime, and use @@ -123,10 +145,34 @@ downloading the models, initialize OpenVINO Runtime, and use and ``*.bin`` files. Then, compile it with ``compile_model()`` to the specified device. +.. code:: ipython3 + + import ipywidgets as widgets + + core = Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + .. code:: ipython3 # Initialize OpenVINO Runtime runtime. - ie_core = Core() + core = Core() def model_init(model_path: str) -> Tuple: @@ -143,16 +189,16 @@ specified device. """ # Read the network and corresponding weights from a file. - model = ie_core.read_model(model=model_path) - # Compile the model for CPU (you can also use GPU). - compiled_model = ie_core.compile_model(model=model, device_name="CPU") + model = core.read_model(model=model_path) + compiled_model = core.compile_model(model=model, device_name=device.value) # Get input and output names of nodes. input_keys = compiled_model.input(0) output_keys = compiled_model.output(0) return input_keys, output_keys, compiled_model -Get attributes from model -~~~~~~~~~~~~~~~~~~~~~~~~~ +Get attributes from model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Use ``input_keys.shape`` to get data shapes. @@ -170,8 +216,9 @@ Use ``input_keys.shape`` to get data shapes. # Get input size - Recognition. height_re, width_re = list(input_key_re.shape)[2:] -Helper function -~~~~~~~~~~~~~~~ +Helper function `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The ``plt_show()`` function is used to show image. @@ -188,8 +235,9 @@ The ``plt_show()`` function is used to show image. plt.axis("off") plt.imshow(raw_image) -Read and display a test image -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Read and display a test image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The input shape of detection model is ``[1, 3, 256, 256]``. Therefore, you need to resize the image to ``256 x 256``, and expand the batch @@ -217,11 +265,12 @@ channel with ``expand_dims`` function. -.. image:: 218-vehicle-detection-and-recognition-with-output_files/218-vehicle-detection-and-recognition-with-output_12_0.png +.. image:: 218-vehicle-detection-and-recognition-with-output_files/218-vehicle-detection-and-recognition-with-output_13_0.png -Use the Detection Model to Detect Vehicles ------------------------------------------- +Use the Detection Model to Detect Vehicles `⇑ <#top>`__ +############################################################################################################################### + .. figure:: https://user-images.githubusercontent.com/47499836/157867076-9e997781-f9ef-45f6-9a51-b515bbf41048.png :alt: pipline @@ -231,14 +280,15 @@ Use the Detection Model to Detect Vehicles As shown in the flowchart, images of individual vehicles are sent to the recognition model. First, use ``infer`` function to get the result. -The detection model output has the format [image_id, label, conf, x_min, -y_min, x_max, y_max], where: +The detection model output has the format +``[image_id, label, conf, x_min, y_min, x_max, y_max]``, where: -- image_id - ID of the image in the batch -- label - predicted class ID (0 - vehicle) -- conf - confidence for the predicted class -- (x_min, y_min) - coordinates of the top left bounding box corner -- (x_max, y_max) - coordinates of the bottom right bounding box corner +- ``image_id`` - ID of the image in the batch +- ``label`` - predicted class ID (0 - vehicle) +- ``conf`` - confidence for the predicted class +- ``(x_min, y_min)`` - coordinates of the top left bounding box corner +- ``(x_max, y_max)`` - coordinates of the bottom right bounding box + corner Delete unused dims and filter out results that are not used. @@ -251,8 +301,9 @@ Delete unused dims and filter out results that are not used. # Remove zero only boxes. boxes = boxes[~np.all(boxes == 0, axis=1)] -Detection Processing -~~~~~~~~~~~~~~~~~~~~ +Detection Processing `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + With the function below, you change the ratio to the real position in the image and filter out low-confidence results. @@ -300,8 +351,9 @@ the image and filter out low-confidence results. # Find the position of a car. car_position = crop_images(image_de, resized_image_de, boxes) -Recognize vehicle attributes -~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Recognize vehicle attributes `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Select one of the detected boxes. Then, crop to an area containing a vehicle to test with the recognition model. Again, you need to resize @@ -320,11 +372,11 @@ the input image and run inference. -.. image:: 218-vehicle-detection-and-recognition-with-output_files/218-vehicle-detection-and-recognition-with-output_19_0.png +.. image:: 218-vehicle-detection-and-recognition-with-output_files/218-vehicle-detection-and-recognition-with-output_20_0.png -Recognition processing -'''''''''''''''''''''' +Recognition processing `⇑ <#top>`__ +----------------------------------------------------------------------------------------------------------------------------------- The result contains colors of the vehicles (white, gray, yellow, red, green, blue, black) and types of vehicles (car, bus, truck, van). Next, @@ -372,8 +424,9 @@ determine the maximum probability as the result. Attributes:('Gray', 'Car') -Combine two models -~~~~~~~~~~~~~~~~~~ +Combine two models `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Congratulations! You successfully used a detection model to crop an image with a vehicle and recognize the attributes of a vehicle. @@ -434,5 +487,5 @@ image with a vehicle and recognize the attributes of a vehicle. -.. image:: 218-vehicle-detection-and-recognition-with-output_files/218-vehicle-detection-and-recognition-with-output_25_0.png +.. image:: 218-vehicle-detection-and-recognition-with-output_files/218-vehicle-detection-and-recognition-with-output_26_0.png diff --git a/docs/notebooks/218-vehicle-detection-and-recognition-with-output_files/218-vehicle-detection-and-recognition-with-output_12_0.png b/docs/notebooks/218-vehicle-detection-and-recognition-with-output_files/218-vehicle-detection-and-recognition-with-output_13_0.png similarity index 100% rename from docs/notebooks/218-vehicle-detection-and-recognition-with-output_files/218-vehicle-detection-and-recognition-with-output_12_0.png rename to docs/notebooks/218-vehicle-detection-and-recognition-with-output_files/218-vehicle-detection-and-recognition-with-output_13_0.png diff --git a/docs/notebooks/218-vehicle-detection-and-recognition-with-output_files/218-vehicle-detection-and-recognition-with-output_19_0.png b/docs/notebooks/218-vehicle-detection-and-recognition-with-output_files/218-vehicle-detection-and-recognition-with-output_20_0.png similarity index 100% rename from docs/notebooks/218-vehicle-detection-and-recognition-with-output_files/218-vehicle-detection-and-recognition-with-output_19_0.png rename to docs/notebooks/218-vehicle-detection-and-recognition-with-output_files/218-vehicle-detection-and-recognition-with-output_20_0.png diff --git a/docs/notebooks/218-vehicle-detection-and-recognition-with-output_files/218-vehicle-detection-and-recognition-with-output_25_0.png b/docs/notebooks/218-vehicle-detection-and-recognition-with-output_files/218-vehicle-detection-and-recognition-with-output_26_0.png similarity index 100% rename from docs/notebooks/218-vehicle-detection-and-recognition-with-output_files/218-vehicle-detection-and-recognition-with-output_25_0.png rename to docs/notebooks/218-vehicle-detection-and-recognition-with-output_files/218-vehicle-detection-and-recognition-with-output_26_0.png diff --git a/docs/notebooks/218-vehicle-detection-and-recognition-with-output_files/index.html b/docs/notebooks/218-vehicle-detection-and-recognition-with-output_files/index.html index 0553ac5c9bb..cdb0d2e6548 100644 --- a/docs/notebooks/218-vehicle-detection-and-recognition-with-output_files/index.html +++ b/docs/notebooks/218-vehicle-detection-and-recognition-with-output_files/index.html @@ -1,9 +1,9 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/218-vehicle-detection-and-recognition-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/218-vehicle-detection-and-recognition-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/218-vehicle-detection-and-recognition-with-output_files/


../
-218-vehicle-detection-and-recognition-with-outp..> 12-Jul-2023 00:11              172680
-218-vehicle-detection-and-recognition-with-outp..> 12-Jul-2023 00:11               19599
-218-vehicle-detection-and-recognition-with-outp..> 12-Jul-2023 00:11              175941
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/218-vehicle-detection-and-recognition-with-output_files/


../
+218-vehicle-detection-and-recognition-with-outp..> 16-Aug-2023 01:31              172680
+218-vehicle-detection-and-recognition-with-outp..> 16-Aug-2023 01:31               19599
+218-vehicle-detection-and-recognition-with-outp..> 16-Aug-2023 01:31              175941
 

diff --git a/docs/notebooks/219-knowledge-graphs-conve-with-output.rst b/docs/notebooks/219-knowledge-graphs-conve-with-output.rst index 628c441e208..bf586695e67 100644 --- a/docs/notebooks/219-knowledge-graphs-conve-with-output.rst +++ b/docs/notebooks/219-knowledge-graphs-conve-with-output.rst @@ -1,21 +1,47 @@ OpenVINO optimizations for Knowledge graphs =========================================== +.. _top: + The goal of this notebook is to showcase performance optimizations for the ConvE knowledge graph embeddings model using the Intel® Distribution of OpenVINO™ Toolkit. The optimizations process contains the following steps: -1. Export the trained model to a format suitable for OpenVINO optimizations and inference -2. Report the inference performance speedup obtained with the optimized OpenVINO model +1. Export the trained model to a format suitable for OpenVINO + optimizations and inference +2. Report the inference performance speedup obtained with the optimized + OpenVINO model -The ConvE model is an implementation of the paper - “Convolutional 2D -Knowledge Graph Embeddings” (https://arxiv.org/abs/1707.01476). The +The ConvE model is an implementation of the paper - +`Convolutional 2D Knowledge Graph Embeddings `__. The sample dataset can be downloaded from: https://github.com/TimDettmers/ConvE/tree/master/countries/countries_S1 -Windows specific settings -------------------------- +**Table of contents**: + +- `Windows specific settings <#windows-specific-settings>`__ +- `Import the packages needed for successful execution <#import-the-packages-needed-for-successful-execution>`__ + + - `Settings: Including path to the serialized model files and input data files <#settings-including-path-to-the-serialized-model-files-and-input-data-files>`__ + - `Download Model Checkpoint <#download-model-checkpoint>`__ + - `Defining the ConvE model class <#defining-the-conve-model-class>`__ + - `Defining the dataloader <#defining-the-dataloader>`__ + - `Evaluate the trained ConvE model <#evaluate-the-trained-conve-model>`__ + - `Prediction on the Knowledge graph. <#prediction-on-the-knowledge-graph>`__ + - `Convert the trained PyTorch model to ONNX format for OpenVINO inference <#convert-the-trained-pytorch-model-to-onnx-format-for-openvino-inference>`__ + - `Evaluate the model performance with OpenVINO <#evaluate-the-model-performance-with-openvino>`__ + +- `Select inference device <#select-inference-device>`__ + + - `Determine the platform specific speedup obtained through OpenVINO graph optimizations <#determine-the-platform-specific-speedup-obtained-through-openvino-graph-optimizations>`__ + - `Benchmark the converted OpenVINO model using benchmark app <#benchmark-the-converted-openvino-model-using-benchmark-app>`__ + - `Conclusions <#conclusions>`__ + - `References <#references>`__ + +Windows specific settings `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -57,8 +83,9 @@ Windows specific settings os.environ["LIB"] = os.pathsep.join(b.library_dirs) print(f"Added {vs_dir} to PATH") -Import the packages needed for successful execution ---------------------------------------------------- +Import the packages needed for successful execution `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -78,8 +105,8 @@ Import the packages needed for successful execution sys.path.append("../utils") from notebook_utils import download_file -Settings: Including path to the serialized model files and input data files -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Settings: Including path to the serialized model files and input data files `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ .. code:: ipython3 @@ -119,8 +146,9 @@ Settings: Including path to the serialized model files and input data files Using cpu device -Download Model Checkpoint -~~~~~~~~~~~~~~~~~~~~~~~~~ +Download Model Checkpoint `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -139,12 +167,13 @@ Download Model Checkpoint .. parsed-literal:: - PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/219-knowledge-graphs-conve/models/conve.pt') + PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/219-knowledge-graphs-conve/models/conve.pt') -Defining the ConvE model class -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Defining the ConvE model class `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -201,8 +230,9 @@ Defining the ConvE model class pred = torch.nn.functional.softmax(x, dim=1) return pred -Defining the dataloader -~~~~~~~~~~~~~~~~~~~~~~~ +Defining the dataloader `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -250,15 +280,15 @@ Defining the dataloader dp.close() return triples_list -Evaluate the trained ConvE model -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Evaluate the trained ConvE model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -First, we will evaluate the model performance using PyTorch. The goal is -to make sure there are no accuracy differences between the original -model inference and the model converted to OpenVINO intermediate -representation inference results. Here, we use a simple accuracy metric -to evaluate the model performance on a test dataset. However, it is -typical to use metrics such as Mean Reciprocal Rank, Hits@10 etc. +First, we will evaluate the model performance using PyTorch. The goal is to make sure there are +no accuracy differences between the original model inference and the +model converted to OpenVINO intermediate representation inference +results. Here, we use a simple accuracy metric to evaluate the model +performance on a test dataset. However, it is typical to use metrics +such as Mean Reciprocal Rank, Hits@10 etc. .. code:: ipython3 @@ -295,19 +325,18 @@ typical to use metrics such as Mean Reciprocal Rank, Hits@10 etc. .. parsed-literal:: - Average time taken for inference: 0.6920695304870605 ms + Average time taken for inference: 0.6897946198781332 ms Mean accuracy of the model on the test dataset: 0.875 -Prediction on the Knowledge graph. -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Prediction on the Knowledge graph. `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Here, we perform the entity prediction on the knowledge graph, as a -sample evaluation task. We pass the source entity ‘san_marino’ and -relation ‘locatedIn’ to the knowledge graph and obtain the target entity -predictions. Expected predictions are target entities that form a -factual triple with the entity and relation passed as inputs to the -knowledge graph. +Here, we perform the entity prediction on the knowledge graph, as a sample evaluation task. +We pass the source entity ``san_marino`` and relation ``locatedIn`` to +the knowledge graph and obtain the target entity predictions. Expected +predictions are target entities that form a factual triple with the +entity and relation passed as inputs to the knowledge graph. .. code:: ipython3 @@ -333,13 +362,13 @@ knowledge graph. Source Entity: san_marino, Relation: locatedin, Target entity prediction: europe -Convert the trained PyTorch model to ONNX format for OpenVINO inference -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Convert the trained PyTorch model to ONNX format for OpenVINO inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -To evaluate performance with OpenVINO, we can either convert the trained -PyTorch model to an intermediate representation (IR) format or to an -ONNX representation. This notebook uses the ONNX format. For more -details on model optimization, refer to: +To evaluate performance with OpenVINO, we can +either convert the trained PyTorch model to an intermediate +representation (IR) format or to an ONNX representation. This notebook +uses the ONNX format. For more details on model optimization, refer to: https://docs.openvino.ai/2023.0/openvino_docs_MO_DG_Deep_Learning_Model_Optimizer_DevGuide.html .. code:: ipython3 @@ -354,8 +383,9 @@ https://docs.openvino.ai/2023.0/openvino_docs_MO_DG_Deep_Learning_Model_Optimize Converting the trained conve model to ONNX format -Evaluate the model performance with OpenVINO -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Evaluate the model performance with OpenVINO `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Now, we evaluate the model performance with the OpenVINO framework. In order to do so, make three main API calls: @@ -369,9 +399,40 @@ Then, the model can be inferred on by using the .. code:: ipython3 - ie = Core() - ir_net = ie.read_model(model=fp32_onnx_path) - compiled_model = ie.compile_model(model=ir_net) + core = Core() + ov_model = core.read_model(model=fp32_onnx_path) + +Select inference device `⇑ <#top>`__ +############################################################################################################################### + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + compiled_model = core.compile_model(model=ov_model, device_name=device.value) input_layer_source = compiled_model.input('input.1') input_layer_relation = compiled_model.input('input.2') output_layer = compiled_model.output(0) @@ -399,12 +460,12 @@ Then, the model can be inferred on by using the .. parsed-literal:: - Average time taken for inference: 1.466284195582072 ms + Average time taken for inference: 1.246631145477295 ms Mean accuracy of the model on the test dataset: 0.10416666666666667 -Determine the platform specific speedup obtained through OpenVINO graph optimizations -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Determine the platform specific speedup obtained through OpenVINO graph optimizations `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ .. code:: ipython3 @@ -413,15 +474,16 @@ Determine the platform specific speedup obtained through OpenVINO graph optimiza .. parsed-literal:: - Speedup with OpenVINO optimizations: 0.47 X + Speedup with OpenVINO optimizations: 0.55 X -Benchmark the converted OpenVINO model using benchmark app -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Benchmark the converted OpenVINO model using benchmark app `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -The OpenVINO toolkit provides a benchmarking application to gauge the -platform specific runtime performance that can be obtained under optimal -configuration parameters for a given model. For more details refer to: +The OpenVINO toolkit provides a benchmarking application to +gauge the platform specific runtime performance that can be obtained +under optimal configuration parameters for a given model. For more +details refer to: https://docs.openvino.ai/2023.0/openvino_inference_engine_tools_benchmark_tool_README.html Here, we use the benchmark application to obtain performance estimates @@ -435,7 +497,7 @@ inference can also be obtained by looking at the benchmark app results. .. code:: ipython3 print('Benchmark OpenVINO model using the benchmark app') - ! benchmark_app -m "$fp32_onnx_path" -d CPU -api async -t 10 -shape "input.1[1],input.2[1]" + ! benchmark_app -m "$fp32_onnx_path" -d device.value -api async -t 10 -shape "input.1[1],input.2[1]" .. parsed-literal:: @@ -445,90 +507,38 @@ inference can also be obtained by looking at the benchmark app results. [ INFO ] Parsing input parameters [Step 2/11] Loading OpenVINO Runtime [ INFO ] OpenVINO: - [ INFO ] Build ................................. 2023.0.0-10926-b4452d56304-releases/2023/0 + [ INFO ] Build ................................. 2023.0.1-11005-fa1c41994f3-releases/2023/0 [ INFO ] [ INFO ] Device info: - [ INFO ] CPU - [ INFO ] Build ................................. 2023.0.0-10926-b4452d56304-releases/2023/0 - [ INFO ] - [ INFO ] - [Step 3/11] Setting device configuration - [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.THROUGHPUT. - [Step 4/11] Reading model files - [ INFO ] Loading model files - [ INFO ] Read model took 20.85 ms - [ INFO ] Original model I/O parameters: - [ INFO ] Model inputs: - [ INFO ] input.1 (node: input.1) : i64 / [...] / [] - [ INFO ] input.2 (node: input.2) : i64 / [...] / [] - [ INFO ] Model outputs: - [ INFO ] 51 (node: 51) : f32 / [...] / [1,271] - [Step 5/11] Resizing model to match image sizes and given batch - [ INFO ] Model batch size: 1 - [ INFO ] Reshaping model: 'input.1': [1], 'input.2': [1] - [ INFO ] Reshape model took 0.97 ms - [Step 6/11] Configuring input of the model - [ INFO ] Model inputs: - [ INFO ] input.1 (node: input.1) : i64 / [...] / [1] - [ INFO ] input.2 (node: input.2) : i64 / [...] / [1] - [ INFO ] Model outputs: - [ INFO ] 51 (node: 51) : f32 / [...] / [1,271] - [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 57.00 ms - [Step 8/11] Querying optimal runtime parameters - [ INFO ] Model: - [ INFO ] NETWORK_NAME: torch_jit - [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 12 - [ INFO ] NUM_STREAMS: 12 - [ INFO ] AFFINITY: Affinity.CORE - [ INFO ] INFERENCE_NUM_THREADS: 24 - [ INFO ] PERF_COUNT: False - [ INFO ] INFERENCE_PRECISION_HINT: - [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT - [ INFO ] EXECUTION_MODE_HINT: ExecutionMode.PERFORMANCE - [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 - [ INFO ] ENABLE_CPU_PINNING: True - [ INFO ] SCHEDULING_CORE_TYPE: SchedulingCoreType.ANY_CORE - [ INFO ] ENABLE_HYPER_THREADING: True - [ INFO ] EXECUTION_DEVICES: ['CPU'] - [Step 9/11] Creating infer requests and preparing input tensors - [ WARNING ] No input files were given for input 'input.1'!. This input will be filled with random values! - [ WARNING ] No input files were given for input 'input.2'!. This input will be filled with random values! - [ INFO ] Fill input 'input.1' with random values - [ INFO ] Fill input 'input.2' with random values - [Step 10/11] Measuring performance (Start inference asynchronously, 12 inference requests, limits: 10000 ms duration) - [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 5.00 ms - [Step 11/11] Dumping statistics report - [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 96480 iterations - [ INFO ] Duration: 10001.93 ms - [ INFO ] Latency: - [ INFO ] Median: 1.01 ms - [ INFO ] Average: 1.04 ms - [ INFO ] Min: 0.60 ms - [ INFO ] Max: 3.45 ms - [ INFO ] Throughput: 9646.14 FPS + [ ERROR ] Check 'false' failed at src/inference/src/core.cpp:84: + Device with "device" name is not registered in the OpenVINO Runtime + Traceback (most recent call last): + File "/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/tools/benchmark/main.py", line 103, in main + benchmark.print_version_info() + File "/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/tools/benchmark/benchmark.py", line 48, in print_version_info + for device, version in self.core.get_versions(self.device).items(): + RuntimeError: Check 'false' failed at src/inference/src/core.cpp:84: + Device with "device" name is not registered in the OpenVINO Runtime + -Conclusions -~~~~~~~~~~~ +Conclusions `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -In this notebook, we convert the trained PyTorch knowledge graph -embeddings model to the OpenVINO format. We confirm that there are no -accuracy differences post conversion. We also perform a sample -evaluation on the knowledge graph. Then, we determine the platform -specific speedup in runtime performance that can be obtained through -OpenVINO graph optimizations. To learn more about the OpenVINO -performance optimizations, refer to: +In this notebook, we convert the trained PyTorch knowledge graph embeddings model to the OpenVINO format. We +confirm that there are no accuracy differences post conversion. We also +perform a sample evaluation on the knowledge graph. Then, we determine +the platform specific speedup in runtime performance that can be +obtained through OpenVINO graph optimizations. To learn more about the +OpenVINO performance optimizations, refer to: https://docs.openvino.ai/2023.0/openvino_docs_optimization_guide_dldt_optimization_guide.html -References -~~~~~~~~~~ +References `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -1. Convolutional 2D Knowledge Graph Embeddings, Tim Dettmers et - al. (https://arxiv.org/abs/1707.01476) -2. Model implementation: https://github.com/TimDettmers/ConvE + 1. Convolutional 2D Knowledge Graph +Embeddings, Tim Dettmers et al. (https://arxiv.org/abs/1707.01476) 2. +Model implementation: https://github.com/TimDettmers/ConvE The ConvE model implementation used in this notebook is licensed under the MIT License. The license is displayed below: MIT License diff --git a/docs/notebooks/220-cross-lingual-books-alignment-with-output.rst b/docs/notebooks/220-cross-lingual-books-alignment-with-output.rst new file mode 100644 index 00000000000..95db7764278 --- /dev/null +++ b/docs/notebooks/220-cross-lingual-books-alignment-with-output.rst @@ -0,0 +1,1075 @@ +Cross-lingual Books Alignment with Transformers and OpenVINO™ +============================================================= + +.. _top: + +Cross-lingual text alignment is the task of matching sentences in a pair +of texts that are translations of each other. In this notebook, you’ll +learn how to use a deep learning model to create a parallel book in +English and German + +This method helps you learn languages but also provides parallel texts +that can be used to train machine translation models. This is +particularly useful if one of the languages is low-resource or you don’t +have enough data to train a full-fledged translation model. + +The notebook shows how to accelerate the most computationally expensive +part of the pipeline - getting vectors from sentences - using the +OpenVINO™ framework. + +Pipeline +-------- + +The notebook guides you through the entire process of creating a +parallel book: from obtaining raw texts to building a visualization of +aligned sentences. Here is the pipeline diagram: + +|image0| + +Visualizing the result allows you to identify areas for improvement in +the pipeline steps, as indicated in the diagram. + +Prerequisites +------------- + +- ``requests`` - for getting books +- ``pysbd`` - for splitting sentences +- ``transformers[torch]`` and ``openvino_dev`` - for getting sentence + embeddings +- ``seaborn`` - for alignment matrix visualization +- ``ipywidgets`` - for displaying HTML and JS output in the notebook + +**Table of contents**: + +- `Get Books <#get-books>`__ +- `Clean Text <#clean-text>`__ +- `Split Text <#split-text>`__ +- `Get Sentence Embeddings <#get-sentence-embeddings>`__ + + - `Optimize the Model with OpenVINO <#optimize-the-model-with-openvino>`__ + +- `Calculate Sentence Alignment <#calculate-sentence-alignment>`__ +- `Postprocess Sentence Alignment <#postprocess-sentence-alignment>`__ +- `Visualize Sentence Alignment <#visualize-sentence-alignment>`__ +- `Speed up Embeddings Computation <#speed-up-embeddings-computation>`__ + +.. |image0| image:: https://user-images.githubusercontent.com/51917466/254582697-18f3ab38-e264-4b2c-a088-8e54b855c1b2.png%22 + +.. code:: ipython3 + + !pip install -q requests pysbd transformers[torch] "openvino_dev>=2023.0" seaborn ipywidgets + + +.. parsed-literal:: + + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + + +Get Books `⇑ <#top>`__ +############################################################################################################################### + + +The first step is to get the books that we will be working with. For +this notebook, we will use English and German versions of Anna Karenina +by Leo Tolstoy. The texts can be obtained from the `Project Gutenberg +site `__. Since copyright laws are complex +and differ from country to country, check the book’s legal availability +in your country. Refer to the Project Gutenberg Permissions, Licensing +and other Common Requests +`page `__ for more +information. + +Find the books on Project Gutenberg `search +page `__ and get the ID of each book. +To get the texts, we will pass the IDs to the +`Gutendex `__ API. + +.. code:: ipython3 + + import requests + + + def get_book_by_id(book_id: int, gutendex_url: str = "https://gutendex.com/") -> str: + book_metadata_url = gutendex_url + "/books/" + str(book_id) + request = requests.get(book_metadata_url, timeout=30) + request.raise_for_status() + + book_metadata = request.json() + book_url = book_metadata["formats"]["text/plain"] + return requests.get(book_url).text + + + en_book_id = 1399 + de_book_id = 44956 + + anna_karenina_en = get_book_by_id(en_book_id) + anna_karenina_de = get_book_by_id(de_book_id) + +Let’s check that we got the right books by showing a part of the texts: + +.. code:: ipython3 + + print(anna_karenina_en[:1500]) + + +.. parsed-literal:: + +  + The Project Gutenberg eBook of Anna Karenina + + This ebook is for the use of anyone anywhere in the United States and + most other parts of the world at no cost and with almost no restrictions + whatsoever. You may copy it, give it away or re-use it under the terms + of the Project Gutenberg License included with this ebook or online + at www.gutenberg.org. If you are not located in the United States, + you will have to check the laws of the country where you are located + before using this eBook. + + + + + Title: Anna Karenina + + Author: graf Leo Tolstoy + Translator: Constance Garnett + + + Release date: July 1, 1998 [eBook #1399]Most recently updated: April 9, 2023 + Language: English + + + + + *** START OF THE PROJECT GUTENBERG EBOOK ANNA KARENINA *** + + [Illustration] + + + + + ANNA KARENINA + + by Leo Tolstoy + + Translated by Constance Garnett + + Contents + + + PART ONE + PART TWO + PART THREE + PART FOUR + PART FIVE + PART SIX + PART SEVEN + PART EIGHT + + + + + PART ONE + + Chapter 1 + + + Happy families are all alike; every unhappy family is unhappy in its + own way. + + Everything was in confusion in the Oblonskys’ house. The wife had + discovered that the husband was carrying on an intrigue with a French + girl, who had been a governess in their family, and she had announced + to her husband that she could not go on living in the same house with + him. This + + +which in a raw format looks like this: + +.. code:: ipython3 + + anna_karenina_en[:1500] + + + + +.. parsed-literal:: + + '\ufeff\r\n The Project Gutenberg eBook of Anna Karenina\r\n \r\nThis ebook is for the use of anyone anywhere in the United States and \r\nmost other parts of the world at no cost and with almost no restrictions \r\nwhatsoever. You may copy it, give it away or re-use it under the terms \r\nof the Project Gutenberg License included with this ebook or online \r\nat www.gutenberg.org. If you are not located in the United States, \r\nyou will have to check the laws of the country where you are located \r\nbefore using this eBook.\r\n\r\n\r\n\r\n \r\n Title: Anna Karenina\r\n \r\n Author: graf Leo Tolstoy\r\n Translator: Constance Garnett\r\n\r\n \r\n Release date: July 1, 1998 [eBook #1399]Most recently updated: April 9, 2023\r\n Language: English\r\n \r\n \r\n \r\n \r\n *** START OF THE PROJECT GUTENBERG EBOOK ANNA KARENINA ***\r\n \r\n[Illustration]\r\n\r\n\r\n\r\n\r\n ANNA KARENINA \r\n\r\n by Leo Tolstoy \r\n\r\n Translated by Constance Garnett \r\n\r\nContents\r\n\r\n\r\n PART ONE\r\n PART TWO\r\n PART THREE\r\n PART FOUR\r\n PART FIVE\r\n PART SIX\r\n PART SEVEN\r\n PART EIGHT\r\n\r\n\r\n\r\n\r\nPART ONE\r\n\r\nChapter 1\r\n\r\n\r\nHappy families are all alike; every unhappy family is unhappy in its\r\nown way.\r\n\r\nEverything was in confusion in the Oblonskys’ house. The wife had\r\ndiscovered that the husband was carrying on an intrigue with a French\r\ngirl, who had been a governess in their family, and she had announced\r\nto her husband that she could not go on living in the same house with\r\nhim. This ' + + + +.. code:: ipython3 + + anna_karenina_de[:1500] + + + + +.. code:: + + '\ufeffThe Project Gutenberg EBook of Anna Karenina, 1. Band, by Leo N. Tolstoi\r\n\r\nThis eBook is for the use of anyone anywhere at no cost and with\r\nalmost no restrictions whatsoever. You may copy it, give it away or\r\nre-use it under the terms of the Project Gutenberg License included\r\nwith this eBook or online at www.gutenberg.org\r\n\r\n\r\nTitle: Anna Karenina, 1. Band\r\n\r\nAuthor: Leo N. Tolstoi\r\n\r\nRelease Date: February 18, 2014 [EBook #44956]\r\n\r\nLanguage: German\r\n\r\n\r\n*** START OF THIS PROJECT GUTENBERG EBOOK ANNA KARENINA, 1. BAND ***\r\n\r\n\r\n\r\n\r\nProduced by Norbert H. Langkau, Jens Nordmann and the\r\nOnline Distributed Proofreading Team at http://www.pgdp.net\r\n\r\n\r\n\r\n\r\n\r\n\r\n\r\n\r\n\r\n Anna Karenina.\r\n\r\n\r\n Roman aus dem Russischen\r\n\r\n des\r\n\r\n Grafen Leo N. Tolstoi.\r\n\r\n\r\n\r\n Nach der siebenten Auflage übersetzt\r\n\r\n von\r\n\r\n Hans Moser.\r\n\r\n\r\n Erster Band.\r\n\r\n\r\n\r\n Leipzig\r\n\r\n Druck und Verlag von Philipp Reclam jun.\r\n\r\n * * * * *\r\n\r\n\r\n\r\n\r\n Erster Teil.\r\n\r\n »Die Rache ist mein, ich will vergelten.«\r\n\r\n 1.\r\n\r\n\r\nAlle glücklichen Familien sind einander ähnlich; jede unglückliche\r\nFamilie ist auf _ihre_ Weise ung' + + + +Clean Text `⇑ <#top>`__ +############################################################################################################################### + + +The downloaded books may contain service information before and after +the main text. The text might have different formatting styles and +markup, for example, phrases from a different language enclosed in +underscores for potential emphasis or italicization: + +.. + + Yes, Alabin was giving a dinner on glass tables, and the tables sang, + *Il mio tesoro— not Il mio tesoro* though, but something better, + and there were some sort of little decanters on the table, and they + were women, too,” he remembered. + +The next stages of the pipeline will be difficult to complete without +cleaning and normalizing the text. Since formatting may differ, manual +work is required at this stage. For example, the main content in the +German version is enclosed in ``* * * * *``, so +it is safe to remove everything before the first occurrence and after +the last occurrence of these asterisks. + +.. hint:: + + There are text-cleaning libraries that clean up common + flaws. If the source of the text is known, you can look for a library + designed for that source, for example + `gutenberg_cleaner `__. + These libraries can reduce manual work and even automate the + process + +.. code:: ipython3 + + import re + from contextlib import contextmanager + from tqdm.auto import tqdm + + + start_pattern_en = r"\nPART ONE" + anna_karenina_en = re.split(start_pattern_en, anna_karenina_en)[1].strip() + + end_pattern_en = "*** END OF THE PROJECT GUTENBERG EBOOK ANNA KARENINA ***" + anna_karenina_en = anna_karenina_en.split(end_pattern_en)[0].strip() + +.. code:: ipython3 + + start_pattern_de = "* * * * *" + anna_karenina_de = anna_karenina_de.split(start_pattern_de, maxsplit=1)[1].strip() + anna_karenina_de = anna_karenina_de.rsplit(start_pattern_de, maxsplit=1)[0].strip() + +.. code:: ipython3 + + anna_karenina_en = anna_karenina_en.replace("\r\n", "\n") + anna_karenina_de = anna_karenina_de.replace("\r\n", "\n") + +For this notebook, we will work only with the first chapter. + +.. code:: ipython3 + + chapter_pattern_en = r"Chapter \d?\d" + chapter_1_en = re.split(chapter_pattern_en, anna_karenina_en)[1].strip() + +.. code:: ipython3 + + chapter_pattern_de = r"\d?\d.\n\n" + chapter_1_de = re.split(chapter_pattern_de, anna_karenina_de)[1].strip() + +Let’s cut it out and define some cleaning functions. + +.. code:: ipython3 + + def remove_single_newline(text: str) -> str: + return re.sub(r"\n(?!\n)", " ", text) + + + def unify_quotes(text: str) -> str: + return re.sub(r"['\"»«“”]", '"', text) + + + def remove_markup(text: str) -> str: + text = text.replace(">=", "").replace("=<", "") + return re.sub(r"_\w|\w_", "", text) + +Combine the cleaning functions into a single pipeline. The ``tqdm`` +library is used to track the execution progress. Define a context +manager to optionally disable the progress indicators if they are not +needed. + +.. code:: ipython3 + + disable_tqdm = False + + + @contextmanager + def disable_tqdm_context(): + global disable_tqdm + disable_tqdm = True + yield + disable_tqdm = False + + + def clean_text(text: str) -> str: + text_cleaning_pipeline = [ + remove_single_newline, + unify_quotes, + remove_markup, + ] + progress_bar = tqdm(text_cleaning_pipeline, disable=disable_tqdm) + for clean_func in progress_bar: + progress_bar.set_postfix_str(clean_func.__name__) + text = clean_func(text) + return text + + + chapter_1_en = clean_text(chapter_1_en) + chapter_1_de = clean_text(chapter_1_de) + + + +.. parsed-literal:: + + 0%| | 0/3 [00:00`__ +############################################################################################################################### + + +Dividing text into sentences is a challenging task in text processing. +The problem is called `sentence boundary +disambiguation `__, +which can be solved using heuristics or machine learning models. This +notebook uses a ``Segmenter`` from the ``pysbd`` library, which is +initialized with an `ISO language +code `__, as the +rules for splitting text into sentences may vary for different +languages. + + **Hint**: The ``book_metadata`` obtained from the Gutendex contains + the language code as well, enabling automation of this part of the + pipeline. + +.. code:: ipython3 + + import pysbd + + + splitter_en = pysbd.Segmenter(language="en", clean=True) + splitter_de = pysbd.Segmenter(language="de", clean=True) + + + sentences_en = splitter_en.segment(chapter_1_en) + sentences_de = splitter_de.segment(chapter_1_de) + + len(sentences_en), len(sentences_de) + + + + +.. parsed-literal:: + + (32, 34) + + + +Get Sentence Embeddings `⇑ <#top>`__ +############################################################################################################################### + + +The next step is to transform sentences into vector representations. +Transformer encoder models, like BERT, provide high-quality embeddings +but can be slow. Additionally, the model should support both chosen +languages. Training separate models for each language pair can be +expensive, so there are many models pre-trained on multiple languages +simultaneously, for example: + +- `multilingual-MiniLM `__ +- `distiluse-base-multilingual-cased `__ +- `bert-base-multilingual-uncased `__ +- `LaBSE `__ + +LaBSE stands for `Language-agnostic BERT Sentence +Embedding `__ and supports 109+ +languages. It has the same architecture as the BERT model but has been +trained on a different task: to produce identical embeddings for +translation pairs. + +|image01| + +This makes LaBSE a great choice for our task and it can be reused for +different language pairs still producing good results. + +.. |image01| image:: https://user-images.githubusercontent.com/51917466/254582913-51531880-373b-40cb-bbf6-1965859df2eb.png%22 + +.. code:: ipython3 + + from typing import List, Union, Dict + from transformers import AutoTokenizer, AutoModel, BertModel + import numpy as np + import torch + from openvino.runtime import CompiledModel as OVModel + + + model_id = "rasa/LaBSE" + pt_model = AutoModel.from_pretrained(model_id) + tokenizer = AutoTokenizer.from_pretrained(model_id) + +The model has two outputs: ``last_hidden_state`` and ``pooler_output``. +For generating embeddings, you can use either the first vector from the +``last_hidden_state``, which corresponds to the special ``[CLS]`` token, +or the entire vector from the second input. Usually, the second option +is used, but we will be using the first option as it also works well for +our task. Fill free to experiment with different outputs to find the +best fit. + +.. code:: ipython3 + + def get_embeddings( + sentences: List[str], + embedding_model: Union[BertModel, OVModel], + ) -> np.ndarray: + if isinstance(embedding_model, OVModel): + embeddings = [ + embedding_model(tokenizer(sent, return_tensors="np").data)[ + "last_hidden_state" + ][0][0] + for sent in tqdm(sentences, disable=disable_tqdm) + ] + return np.vstack(embeddings) + else: + embeddings = [ + embedding_model(**tokenizer(sent, return_tensors="pt"))[ + "last_hidden_state" + ][0][0] + for sent in tqdm(sentences, disable=disable_tqdm) + ] + return torch.vstack(embeddings) + + + embeddings_en_pt = get_embeddings(sentences_en, pt_model) + embeddings_de_pt = get_embeddings(sentences_de, pt_model) + + + +.. parsed-literal:: + + 0%| | 0/32 [00:00`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +The LaBSE model is quite large and can be slow to infer on some +hardware, so let’s optimize it with OpenVINO. `Model conversion Python +API `__ +accepts the PyTorch/Transformers model object and additional information +about model inputs. An ``example_input`` is needed to trace the model +execution graph, as PyTorch constructs it dynamically during inference. +The converted model must be compiled for the target device using the +``Core`` object before it can be used. + +.. code:: ipython3 + + from openvino.runtime import Core, Type + from openvino.tools.mo import convert_model + + + # 3 inputs with dynamic axis [batch_size, sequence_length] and type int64 + inputs_info = [([-1, -1], Type.i64)] * 3 + ov_model = convert_model( + pt_model, + example_input=tokenizer("test", return_tensors="pt").data, + input=inputs_info, + ) + + core = Core() + compiled_model = core.compile_model(ov_model, "CPU") + + embeddings_en = get_embeddings(sentences_en, compiled_model) + embeddings_de = get_embeddings(sentences_de, compiled_model) + + +.. parsed-literal:: + + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/jit/annotations.py:309: UserWarning: TorchScript will treat type annotations of Tensor dtype-specific subtypes as if they are normal Tensors. dtype constraints are not enforced in compilation either. + warnings.warn("TorchScript will treat type annotations of Tensor " + + +.. parsed-literal:: + + huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks... + To disable this warning, you can either: + - Avoid using `tokenizers` before the fork if possible + - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false) + + + +.. parsed-literal:: + + 0%| | 0/32 [00:00`__ +############################################################################################################################### + + +With the embedding matrices from the previous step, we can calculate the +alignment: 1. Calculate sentence similarity between each pair of +sentences. 1. Transform the values in the similarity matrix rows and +columns to a specified range, for example ``[-1, 1]``. 1. Compare the +values with a threshold to get boolean matrices with 0 and 1. 1. +Sentence pairs that have 1 in both matrices should be aligned according +to the model. + +We visualize the resulting matrix and also make sure that the result of +the converted model is the same as the original one. + +.. code:: ipython3 + + import seaborn as sns + import matplotlib.pyplot as plt + + + sns.set_style("whitegrid") + + + def transform(x): + x = x - np.mean(x) + return x / np.var(x) + + + def calculate_alignment_matrix( + first: np.ndarray, second: np.ndarray, threshold: float = 1e-3 + ) -> np.ndarray: + similarity = first @ second.T # 1 + similarity_en_to_de = np.apply_along_axis(transform, -1, similarity) # 2 + similarity_de_to_en = np.apply_along_axis(transform, -2, similarity) # 2 + + both_one = (similarity_en_to_de > threshold) * ( + similarity_de_to_en > threshold + ) # 3 and 4 + return both_one + + + threshold = 0.028 + + alignment_matrix = calculate_alignment_matrix(embeddings_en, embeddings_de, threshold) + alignment_matrix_pt = calculate_alignment_matrix( + embeddings_en_pt.detach().numpy(), + embeddings_de_pt.detach().numpy(), + threshold, + ) + + graph, axis = plt.subplots(1, 2, figsize=(10, 5), sharey=True) + + for matrix, ax, title in zip( + (alignment_matrix, alignment_matrix_pt), axis, ("OpenVINO", "PyTorch") + ): + plot = sns.heatmap(matrix, cbar=False, square=True, ax=ax) + plot.set_title(f"Sentence Alignment Matrix {title}") + plot.set_xlabel("German") + if title == "OpenVINO": + plot.set_ylabel("English") + + graph.tight_layout() + + + +.. image:: 220-cross-lingual-books-alignment-with-output_files/220-cross-lingual-books-alignment-with-output_31_0.png + + +After visualizing and comparing the alignment matrices, let’s transform +them into a dictionary to make it more convenient to work with alignment +in Python. Dictionary keys will be English sentence numbers and values +will be lists of German sentence numbers. + +.. code:: ipython3 + + def make_alignment(alignment_matrix: np.ndarray) -> Dict[int, List[int]]: + aligned = {idx: [] for idx, sent in enumerate(sentences_en)} + for en_idx, de_idx in zip(*np.nonzero(alignment_matrix)): + aligned[en_idx].append(de_idx) + return aligned + + + aligned = make_alignment(alignment_matrix) + aligned + + + + +.. parsed-literal:: + + {0: [0], + 1: [2], + 2: [3], + 3: [4], + 4: [5], + 5: [6], + 6: [7], + 7: [8], + 8: [9, 10], + 9: [11], + 10: [13, 14], + 11: [15], + 12: [16], + 13: [17], + 14: [], + 15: [18], + 16: [19], + 17: [20], + 18: [21], + 19: [23], + 20: [24], + 21: [25], + 22: [26], + 23: [], + 24: [27], + 25: [], + 26: [28], + 27: [29], + 28: [30], + 29: [31], + 30: [32], + 31: [33]} + + + +Postprocess Sentence Alignment `⇑ <#top>`__ +############################################################################################################################### + + +There are several gaps in the resulting alignment, such as English +sentence #14 not mapping to any German sentence. Here are some possible +reasons for this: + +1. There are no equivalent sentences in the other book, and in such + cases, the model is working correctly. +2. The sentence has an equivalent sentence in another language, but the + model failed to identify it. The ``threshold`` might be too high, or + the model is not sensitive enough. To address this, lower the + ``threshold`` value or try a different model. +3. The sentence has an equivalent text part in another language, meaning + that either the sentence splitters are too fine or too coarse. Try + tuning the text cleaning and splitting steps to fix this issue. +4. Combination of 2 and 3, where both the model’s sensitivity and text + preparation steps need adjustments. + +Another solution to address this issue is by applying heuristics. As you +can see, English sentence 13 corresponds to German 17, and 15 to 18. +Most likely, English sentence 14 is part of either German sentence 17 or +18. By comparing the similarity using the model, you can choose the most +suitable alignment. + +Visualize Sentence Alignment `⇑ <#top>`__ +############################################################################################################################### + + +To evaluate the final alignment and choose the best way to improve the +results of the pipeline, we will create an interactive table with HTML +and JS. + +.. code:: ipython3 + + from IPython.display import display, HTML + from itertools import zip_longest + from io import StringIO + + + def create_interactive_table( + list1: List[str], list2: List[str], mapping: Dict[int, List[int]] + ) -> str: + def inverse_mapping(mapping): + inverse_map = {idx: [] for idx in range(len(list2))} + + for key, values in mapping.items(): + for value in values: + inverse_map[value].append(key) + + return inverse_map + + inversed_mapping = inverse_mapping(mapping) + + table_html = StringIO() + table_html.write( + '' + ) + for i, (first, second) in enumerate(zip_longest(list1, list2)): + table_html.write("") + if i < len(list1): + table_html.write(f'') + else: + table_html.write("") + if i < len(list2): + table_html.write(f'') + else: + table_html.write("") + table_html.write("") + + table_html.write("
Sentences ENSentences DE
{first}{second}
") + + hover_script = ( + """ + + """ + ) + table_html.write(hover_script) + return table_html.getvalue() + +.. code:: ipython3 + + html_code = create_interactive_table(sentences_en, sentences_de, aligned) + display(HTML(html_code)) + + + +.. raw:: html + +
Sentences ENSentences DE
Happy families are all alike; every unhappy family is unhappy in its own way.Alle glücklichen Familien sind einander ähnlich; jede unglückliche Familie ist auf hr Weise unglücklich.
Everything was in confusion in the Oblonskys’ house.--
The wife had discovered that the husband was carrying on an intrigue with a French girl, who had been a governess in their family, and she had announced to her husband that she could not go on living in the same house with him.Im Hause der Oblonskiy herrschte allgemeine Verwirrung.
This position of affairs had now lasted three days, and not only the husband and wife themselves, but all the members of their family and household, were painfully conscious of it.Die Dame des Hauses hatte in Erfahrung gebracht, daß ihr Gatte mit der im Hause gewesenen französischen Gouvernante ein Verhältnis unterhalten, und ihm erklärt, sie könne fürderhin nicht mehr mit ihm unter einem Dache bleiben.
Every person in the house felt that there was no sense in their living together, and that the stray people brought together by chance in any inn had more in common with one another than they, the members of the family and household of the Oblonskys.Diese Situation währte bereits seit drei Tagen und sie wurde nicht allein von den beiden Ehegatten selbst, nein auch von allen Familienmitgliedern und dem Personal aufs Peinlichste empfunden.
The wife did not leave her own room, the husband had not been at home for three days.Sie alle fühlten, daß in ihrem Zusammenleben kein höherer Gedanke mehr liege, daß die Leute, welche auf jeder Poststation sich zufällig träfen, noch enger zu einander gehörten, als sie, die Glieder der Familie selbst, und das im Hause geborene und aufgewachsene Gesinde der Oblonskiy.
The children ran wild all over the house; the English governess quarreled with the housekeeper, and wrote to a friend asking her to look out for a new situation for her; the man-cook had walked off the day before just at dinner time; the kitchen-maid, and the coachman had given warning.Die Herrin des Hauses verließ ihre Gemächer nicht, der Gebieter war schon seit drei Tagen abwesend.
Three days after the quarrel, Prince Stepan Arkadyevitch Oblonsky—Stiva, as he was called in the fashionable world—woke up at his usual hour, that is, at eight o’clock in the morning, not in his wife’s bedroom, but on the leather-covered sofa in his study.Die Kinder liefen wie verwaist im ganzen Hause umher, die Engländerin schalt auf die Wirtschafterin und schrieb an eine Freundin, diese möchte ihr eine neue Stellung verschaffen, der Koch hatte bereits seit gestern um die Mittagszeit das Haus verlassen und die Köchin, sowie der Kutscher hatten ihre Rechnungen eingereicht.
He turned over his stout, well-cared-for person on the springy sofa, as though he would sink into a long sleep again; he vigorously embraced the pillow on the other side and buried his face in it; but all at once he jumped up, sat up on the sofa, and opened his eyes.Am dritten Tage nach der Scene erwachte der Fürst Stefan Arkadjewitsch Oblonskiy -- Stiwa hieß er in der Welt -- um die gewöhnliche Stunde, das heißt um acht Uhr morgens, aber nicht im Schlafzimmer seiner Gattin, sondern in seinem Kabinett auf dem Saffiandiwan.
"Yes, yes, how was it now?" he thought, going over his dream.Er wandte seinen vollen verweichlichten Leib auf den Sprungfedern des Diwans, als wünsche er noch weiter zu schlafen, während er von der andern Seite innig ein Kissen umfaßte und an die Wange drückte.
"Now, how was it? To be sure! Alabin was giving a dinner at Darmstadt; no, not Darmstadt, but something American. Yes, but then, Darmstadt was in America. Yes, Alabin was giving a dinner on glass tables, and the tables sang, l mio tesor—not l mio tesor though, but something better, and there were some sort of little decanters on the table, and they were women, too," he remembered.Plötzlich aber sprang er empor, setzte sich aufrecht und öffnete die Augen.
Stepan Arkadyevitch’s eyes twinkled gaily, and he pondered with a smile."Ja, ja, wie war doch das?" sann er, über seinem Traum grübelnd.
"Yes, it was nice, very nice. There was a great deal more that was delightful, only there’s no putting it into words, or even expressing it in one’s thoughts awake.""Wie war doch das?
And noticing a gleam of light peeping in beside one of the serge curtains, he cheerfully dropped his feet over the edge of the sofa, and felt about with them for his slippers, a present on his last birthday, worked for him by his wife on gold-colored morocco.Richtig; Alabin gab ein Diner in Darmstadt; nein, nicht in Darmstadt, es war so etwas Amerikanisches dabei.
And, as he had done every day for the last nine years, he stretched out his hand, without getting up, towards the place where his dressing-gown always hung in his bedroom.Dieses Darmstadt war aber in Amerika, ja, und Alabin gab das Essen auf gläsernen Tischen, ja, und die Tische sangen: Il mio tesoro -- oder nicht so, es war etwas Besseres, und gewisse kleine Karaffen, wie Frauenzimmer aussehend," -- fiel ihm ein.
And thereupon he suddenly remembered that he was not sleeping in his wife’s room, but in his study, and why: the smile vanished from his face, he knitted his brows.Die Augen Stefan Arkadjewitschs blitzten heiter, er sann und lächelte.
"Ah, ah, ah! Oo!..." he muttered, recalling everything that had happened."Ja, es war hübsch, sehr hübsch. Es gab viel Ausgezeichnetes dabei, was man mit Worten nicht schildern könnte und in Gedanken nicht ausdrücken."
And again every detail of his quarrel with his wife was present to his imagination, all the hopelessness of his position, and worst of all, his own fault.Er bemerkte einen Lichtstreif, der sich von der Seite durch die baumwollenen Stores gestohlen hatte und schnellte lustig mit den Füßen vom Sofa, um mit ihnen die von seiner Gattin ihm im vorigen Jahr zum Geburtstag verehrten gold- und saffiangestickten Pantoffeln zu suchen; während er, einer alten neunjährigen Gewohnheit folgend, ohne aufzustehen mit der Hand nach der Stelle fuhr, wo in dem Schlafzimmer sonst sein Morgenrock zu hängen pflegte.
"Yes, she won’t forgive me, and she can’t forgive me. And the most awful thing about it is that it’s all my fault—all my fault, though I’m not to blame. That’s the point of the whole situation," he reflected.Hierbei erst kam er zur Besinnung; er entsann sich jäh wie es kam, daß er nicht im Schlafgemach seiner Gattin, sondern in dem Kabinett schlief; das Lächeln verschwand von seinen Zügen und er runzelte die Stirn.
"Oh, oh, oh!" he kept repeating in despair, as he remembered the acutely painful sensations caused him by this quarrel."O, o, o, ach," brach er jammernd aus, indem ihm alles wieder einfiel, was vorgefallen war.
Most unpleasant of all was the first minute when, on coming, happy and good-humored, from the theater, with a huge pear in his hand for his wife, he had not found his wife in the drawing-room, to his surprise had not found her in the study either, and saw her at last in her bedroom with the unlucky letter that revealed everything in her hand.Vor seinem Innern erstanden von neuem alle die Einzelheiten des Auftritts mit seiner Frau, erstand die ganze Mißlichkeit seiner Lage und -- was ihm am peinlichsten war -- seine eigene Schuld.
She, his Dolly, forever fussing and worrying over household details, and limited in her ideas, as he considered, was sitting perfectly still with the letter in her hand, looking at him with an expression of horror, despair, and indignation."Ja wohl, sie wird nicht verzeihen, sie kann nicht verzeihen, und am Schrecklichsten ist, daß die Schuld an allem nur ich selbst trage -- ich bin schuld -- aber nicht schuldig!
"What’s this? this?" she asked, pointing to the letter.Und hierin liegt das ganze Drama," dachte er, "o weh, o weh!"
And at this recollection, Stepan Arkadyevitch, as is so often the case, was not so much annoyed at the fact itself as at the way in which he had met his wife’s words.Er sprach voller Verzweiflung, indem er sich alle die tiefen Eindrücke vergegenwärtigte, die er in jener Scene erhalten.
There happened to him at that instant what does happen to people when they are unexpectedly caught in something very disgraceful.Am unerquicklichsten war ihm jene erste Minute gewesen, da er, heiter und zufrieden aus dem Theater heimkehrend, eine ungeheure Birne für seine Frau in der Hand, diese weder im Salon noch im Kabinett fand, und sie endlich im Schlafzimmer antraf, jenen unglückseligen Brief, der alles entdeckte, in den Händen.
He did not succeed in adapting his face to the position in which he was placed towards his wife by the discovery of his fault.Sie, die er für die ewig sorgende, ewig sich mühende, allgegenwärtige Dolly gehalten, sie saß jetzt regungslos, den Brief in der Hand, mit dem Ausdruck des Entsetzens, der Verzweiflung und der Wut ihm entgegenblickend.
Instead of being hurt, denying, defending himself, begging forgiveness, instead of remaining indifferent even—anything would have been better than what he did do—his face utterly involuntarily (reflex spinal action, reflected Stepan Arkadyevitch, who was fond of physiology)—utterly involuntarily assumed its habitual, good-humored, and therefore idiotic smile."Was ist das?" frug sie ihn, auf das Schreiben weisend, und in der Erinnerung hieran quälte ihn, wie das oft zu geschehen pflegt, nicht sowohl der Vorfall selbst, als die Art, wie er ihr auf diese Worte geantwortet hatte.
This idiotic smile he could not forgive himself.Es ging ihm in diesem Augenblick, wie den meisten Menschen, wenn sie unerwartet eines zu schmählichen Vergehens überführt werden.
Catching sight of that smile, Dolly shuddered as though at physical pain, broke out with her characteristic heat into a flood of cruel words, and rushed out of the room.Er verstand nicht, sein Gesicht der Situation anzupassen, in welche er nach der Entdeckung seiner Schuld geraten war, und anstatt den Gekränkten zu spielen, sich zu verteidigen, sich zu rechtfertigen und um Verzeihung zu bitten oder wenigstens gleichmütig zu bleiben -- alles dies wäre noch besser gewesen als das, was er wirklich that -- verzogen sich seine Mienen ("Gehirnreflexe" dachte Stefan Arkadjewitsch, als Liebhaber von Physiologie) unwillkürlich und plötzlich zu seinem gewohnten, gutmütigen und daher ziemlich einfältigen Lächeln.
Since then she had refused to see her husband.Dieses dumme Lächeln konnte er sich selbst nicht vergeben.
"It’s that idiotic smile that’s to blame for it all," thought Stepan Arkadyevitch.Als Dolly es gewahrt hatte, erbebte sie, wie von einem physischen Schmerz, und erging sich dann mit der ihr eigenen Leidenschaftlichkeit in einem Strom bitterer Worte, worauf sie das Gemach verließ.
"But what’s to be done? What’s to be done?" he said to himself in despair, and found no answer.Von dieser Zeit an wollte sie ihren Gatten nicht mehr sehen.
"An allem ist das dumme Lächeln schuld," dachte Stefan Arkadjewitsch.
"Aber was soll ich thun, was soll ich thun?" frug er voll Verzweiflung sich selbst, ohne eine Antwort zu finden.
+ + + + +You can see that the pipeline does not fully clean up the German text, +resulting in issues like the second sentence consisting of only ``--``. +On a positive note, the split sentences in the German translation line +up correctly with the single English sentence. Overall, the pipeline +already works well, but there is still room for improvement. + +Save the OpenVINO model to disk for future use: + +.. code:: ipython3 + + from openvino.runtime import serialize + + + ov_model_path = "ov_model/model.xml" + serialize(ov_model, ov_model_path) + +To read the model from disk, use the ``read_model`` method of the +``Core`` object: + +.. code:: ipython3 + + ov_model = core.read_model(ov_model_path) + +Speed up Embeddings Computation `⇑ <#top>`__ +############################################################################################################################### + + +Let’s see how we can speed up the most computationally complex part of +the pipeline - getting embeddings. You might wonder why, when using +OpenVINO, you need to compile the model after reading it. There are two +main reasons for this: 1. Compatibility with different devices. The +model can be compiled to run on a `specific +device `__, +like CPU, GPU or GNA. Each device may work with different data types, +support different features, and gain performance by changing the neural +network for a specific computing model. With OpenVINO, you do not need +to store multiple copies of the network with optimized for different +hardware. A universal OpenVINO model representation is enough. 1. +Optimization for different scenarios. For example, one scenario +prioritizes minimizing the *time between starting and finishing model +inference* (`latency-oriented +optimization `__). +In our case, it is more important *how many texts per second the model +can process* (`throughput-oriented +optimization `__). + +To get a throughput-optimized model, pass a `performance +hint `__ +as a configuration during compilation. Then OpenVINO selects the optimal +parameters for execution on the available hardware. + +.. code:: ipython3 + + from openvino.runtime import Core, AsyncInferQueue, InferRequest + from typing import Any + + + compiled_throughput_hint = core.compile_model( + ov_model, + device_name="CPU", + config={"PERFORMANCE_HINT": "THROUGHPUT"}, + ) + +To further optimize hardware utilization, let’s change the inference +mode from synchronous (Sync) to asynchronous (Async). While the +synchronous API may be easier to start with, it is +`recommended `__ +to use the asynchronous (callbacks-based) API in production code. It is +the most general and scalable way to implement flow control for any +number of requests. + +To work in asynchronous mode, you need to define two things: + +1. Instantiate an ``AsyncInferQueue``, which can then be populated with + inference requests. +2. Define a ``callback`` function that will be called after the output + request has been executed and its results have been processed. + +In addition to the model input, any data required for post-processing +can be passed to the queue. We can create a zero embedding matrix in +advance and fill it in as the inference requests are executed. + +.. code:: ipython3 + + def get_embeddings_async(sentences: List[str], embedding_model: OVModel) -> np.ndarray: + def callback(infer_request: InferRequest, user_data: List[Any]) -> None: + embeddings, idx, pbar = user_data + embedding = infer_request.get_output_tensor(0).data[0, 0] + embeddings[idx] = embedding + pbar.update() + + infer_queue = AsyncInferQueue(embedding_model) + infer_queue.set_callback(callback) + + embedding_dim = ( + embedding_model.output(0).get_partial_shape().get_dimension(2).get_length() + ) + embeddings = np.zeros((len(sentences), embedding_dim)) + + with tqdm(total=len(sentences), disable=disable_tqdm) as pbar: + for idx, sent in enumerate(sentences): + tokenized = tokenizer(sent, return_tensors="np").data + + infer_queue.start_async(tokenized, [embeddings, idx, pbar]) + + infer_queue.wait_all() + + return embeddings + +Let’s compare the models and plot the results. + + **Note**: To get a more accurate benchmark, use the `Benchmark Python + Tool `__ + +.. code:: ipython3 + + number_of_chars = 15_000 + more_sentences_en = splitter_en.segment(clean_text(anna_karenina_en[:number_of_chars])) + len(more_sentences_en) + + + +.. parsed-literal:: + + 0%| | 0/3 [00:00`__ +- `OpenVINO Async +API `__ +- `Throughput +Optimizations `__ diff --git a/docs/notebooks/220-cross-lingual-books-alignment-with-output_files/220-cross-lingual-books-alignment-with-output_31_0.png b/docs/notebooks/220-cross-lingual-books-alignment-with-output_files/220-cross-lingual-books-alignment-with-output_31_0.png new file mode 100644 index 00000000000..153bcc7bee1 --- /dev/null +++ b/docs/notebooks/220-cross-lingual-books-alignment-with-output_files/220-cross-lingual-books-alignment-with-output_31_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:df7720db03699025cc9a4f4a6c63b5a062961e1051ba2f8183328233097dce7e +size 24464 diff --git a/docs/notebooks/220-cross-lingual-books-alignment-with-output_files/220-cross-lingual-books-alignment-with-output_48_0.png b/docs/notebooks/220-cross-lingual-books-alignment-with-output_files/220-cross-lingual-books-alignment-with-output_48_0.png new file mode 100644 index 00000000000..a9f800711b1 --- /dev/null +++ b/docs/notebooks/220-cross-lingual-books-alignment-with-output_files/220-cross-lingual-books-alignment-with-output_48_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:2a4a7fe4908f5932347d616812fa74b28587eb644058a6f19a03dbab178a31b4 +size 31382 diff --git a/docs/notebooks/220-cross-lingual-books-alignment-with-output_files/index.html b/docs/notebooks/220-cross-lingual-books-alignment-with-output_files/index.html new file mode 100644 index 00000000000..fbb2fd2341b --- /dev/null +++ b/docs/notebooks/220-cross-lingual-books-alignment-with-output_files/index.html @@ -0,0 +1,8 @@ + +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/220-cross-lingual-books-alignment-with-output_files/ + +

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/220-cross-lingual-books-alignment-with-output_files/


../
+220-cross-lingual-books-alignment-with-output_3..> 16-Aug-2023 01:31               24464
+220-cross-lingual-books-alignment-with-output_4..> 16-Aug-2023 01:31               31382
+

+ diff --git a/docs/notebooks/221-machine-translation-with-output.rst b/docs/notebooks/221-machine-translation-with-output.rst index bf915c9e6cf..f8c36d8b482 100644 --- a/docs/notebooks/221-machine-translation-with-output.rst +++ b/docs/notebooks/221-machine-translation-with-output.rst @@ -1,11 +1,13 @@ Machine translation demo ======================== +.. _top: + This demo utilizes Intel’s pre-trained model that translates from English to German. More information about the model can be found `here `__. -This model encodes sentences using the SentecePieceBPETokenizer from +This model encodes sentences using the ``SentecePieceBPETokenizer`` from HuggingFace. The tokenizer vocabulary is downloaded automatically with the OMZ tool. @@ -13,15 +15,33 @@ the OMZ tool. following structure: ```` + *tokenized sentence* + ```` + ```` (```` tokens pad the remaining blank spaces). -**Ouput** After the inference, we have a sequence of up to 200 tokens. -The structure is the same as the one for the input. +**Output** After the inference, we have a sequence of up to 200 tokens. +The structure is the same as the one for the input. + +**Table of contents**: + +- `Downloading model <#downloading-model>`__ +- `Load and configure the model <#load-and-configure-the-model>`__ +- `Select inference device <#select-inference-device>`__ +- `Load tokenizers <#load-tokenizers>`__ +- `Perform translation <#perform-translation>`__ +- `Translate the sentence <#translate-the-sentence>`__ + + - `Test your translation <#test-your-translation>`__ .. code:: ipython3 # Install requirements - !pip install -q 'openvino-dev>=2023.0.0' + !pip install -q "openvino-dev>=2023.0.0" !pip install -q tokenizers + +.. parsed-literal:: + + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + + .. code:: ipython3 import time @@ -30,11 +50,11 @@ The structure is the same as the one for the input. import itertools from tokenizers import SentencePieceBPETokenizer -Downloading model ------------------ +Downloading model `⇑ <#top>`__ +############################################################################################################################### -The following command will download the model to the current directory. -Make sure you have run ``pip install openvino-dev`` beforehand. +The following command will download the model to the current directory. Make sure you have run +``pip install openvino-dev`` beforehand. .. code:: ipython3 @@ -45,55 +65,89 @@ Make sure you have run ``pip install openvino-dev`` beforehand. ################|| Downloading machine-translation-nar-en-de-0002 ||################ - ========== Downloading /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/221-machine-translation/intel/machine-translation-nar-en-de-0002/tokenizer_tgt/merges.txt + ========== Downloading /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/221-machine-translation/intel/machine-translation-nar-en-de-0002/tokenizer_tgt/merges.txt - ========== Downloading /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/221-machine-translation/intel/machine-translation-nar-en-de-0002/tokenizer_tgt/vocab.json + ========== Downloading /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/221-machine-translation/intel/machine-translation-nar-en-de-0002/tokenizer_tgt/vocab.json - ========== Downloading /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/221-machine-translation/intel/machine-translation-nar-en-de-0002/tokenizer_src/merges.txt + ========== Downloading /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/221-machine-translation/intel/machine-translation-nar-en-de-0002/tokenizer_src/merges.txt - ========== Downloading /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/221-machine-translation/intel/machine-translation-nar-en-de-0002/tokenizer_src/vocab.json + ========== Downloading /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/221-machine-translation/intel/machine-translation-nar-en-de-0002/tokenizer_src/vocab.json - ========== Downloading /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/221-machine-translation/intel/machine-translation-nar-en-de-0002/FP32/machine-translation-nar-en-de-0002.xml + ========== Downloading /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/221-machine-translation/intel/machine-translation-nar-en-de-0002/FP32/machine-translation-nar-en-de-0002.xml - ========== Downloading /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/221-machine-translation/intel/machine-translation-nar-en-de-0002/FP32/machine-translation-nar-en-de-0002.bin + ========== Downloading /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/221-machine-translation/intel/machine-translation-nar-en-de-0002/FP32/machine-translation-nar-en-de-0002.bin - ========== Downloading /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/221-machine-translation/intel/machine-translation-nar-en-de-0002/FP16/machine-translation-nar-en-de-0002.xml + ========== Downloading /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/221-machine-translation/intel/machine-translation-nar-en-de-0002/FP16/machine-translation-nar-en-de-0002.xml - ========== Downloading /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/221-machine-translation/intel/machine-translation-nar-en-de-0002/FP16/machine-translation-nar-en-de-0002.bin + ========== Downloading /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/221-machine-translation/intel/machine-translation-nar-en-de-0002/FP16/machine-translation-nar-en-de-0002.bin -Load and configure the model ----------------------------- +Load and configure the model `⇑ <#top>`__ +############################################################################################################################### -The model is now available in the ``intel/`` folder. Below, we load and -configure its inputs and outputs. +The model is now available in the ``intel/`` folder. Below, we load and configure its inputs and +outputs. .. code:: ipython3 core = Core() model = core.read_model('intel/machine-translation-nar-en-de-0002/FP32/machine-translation-nar-en-de-0002.xml') - compiled_model = core.compile_model(model) input_name = "tokens" output_name = "pred" model.output(output_name) max_tokens = model.input(input_name).shape[1] -Load tokenizers ---------------- +Select inference device `⇑ <#top>`__ +############################################################################################################################### + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + core = Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + compiled_model = core.compile_model(model, device.value) + +Load tokenizers `⇑ <#top>`__ +############################################################################################################################### + NLP models usually take a list of tokens as standard input. A token is a single word converted to some integer. To provide the proper input, we -need the vocabulary for such mapping. We use ‘merges.txt’ to find out -what sequences of letters form a token. ‘vocab.json’ specifies the +need the vocabulary for such mapping. We use ``merges.txt`` to find out +what sequences of letters form a token. ``vocab.json`` specifies the mapping between tokens and integers. The input needs to be transformed into a token sequence the model @@ -114,8 +168,8 @@ Initialize the tokenizer for the input ``src_tokenizer`` and the output 'intel/machine-translation-nar-en-de-0002/tokenizer_tgt/merges.txt' ) -Perform translation -------------------- +Perform translation `⇑ <#top>`__ +############################################################################################################################### The following function translates a sentence in English to German. @@ -163,8 +217,8 @@ The following function translates a sentence in English to German. sentence = " ".join(key for key, _ in itertools.groupby(sentence)) return sentence -Translate the sentence ----------------------- +Translate the sentence `⇑ <#top>`__ +############################################################################################################################### The following function is a basic loop that translates sentences. @@ -193,11 +247,10 @@ The following function is a basic loop that translates sentences. # uncomment the following line for a real time translation of your input # run_translator() -Test your translation -~~~~~~~~~~~~~~~~~~~~~ +Test your translation `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Run the following cell with an English sentence to have it translated to -German +Run the following cell with an English sentence to have it translated to German .. code:: ipython3 diff --git a/docs/notebooks/222-vision-image-colorization-with-output.rst b/docs/notebooks/222-vision-image-colorization-with-output.rst index c7cc0680c04..6117364879d 100644 --- a/docs/notebooks/222-vision-image-colorization-with-output.rst +++ b/docs/notebooks/222-vision-image-colorization-with-output.rst @@ -1,6 +1,8 @@ Image Colorization with OpenVINO ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +.. _top: + This notebook demonstrates how to colorize images with OpenVINO using the Colorization model `colorization-v2 `__ @@ -40,10 +42,25 @@ About Colorization-siggraph A- and B-channels of LAB-image as output. See the `colorization `__ -repository for more details. +repository for more details. + +**Table of contents**: + +- `Imports <#imports>`__ +- `Configurations <#configurations>`__ + + - `Select inference device <#select-inference-device>`__ + +- `Download the model <#download-the-model>`__ +- `Convert the model to OpenVINO IR <#convert-the-model-to-openvino-ir>`__ +- `Loading the Model <#loading-the-model>`__ +- `Utility Functions <#utility-functions>`__ +- `Load the Image <#load-the-image>`__ +- `Display Colorized Image <#display-colorized-image>`__ + +Imports `⇑ <#top>`__ +############################################################################################################################### -Imports -------- .. code:: ipython3 @@ -59,8 +76,9 @@ Imports sys.path.append("../utils") import notebook_utils as utils -Configurations --------------- +Configurations `⇑ <#top>`__ +############################################################################################################################### + - ``PRECISION`` - {FP16, FP32}, default: FP16. - ``MODEL_DIR`` - directory where the model is to be stored, default: @@ -68,7 +86,6 @@ Configurations - ``MODEL_NAME`` - name of the model used for inference, default: colorization-v2. - ``DATA_DIR`` - directory where test images are stored, default: data. -- ``DEVICE`` - {CPU, GPU} device to used for inference, default: CPU. .. code:: ipython3 @@ -78,10 +95,40 @@ Configurations # MODEL_NAME="colorization-siggraph" MODEL_PATH = f"{MODEL_DIR}/public/{MODEL_NAME}/{PRECISION}/{MODEL_NAME}.xml" DATA_DIR = "data" - DEVICE = "CPU" -Download the model ------------------- +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + core = Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +Download the model `⇑ <#top>`__ +############################################################################################################################### + ``omz_downloader`` downloads model files from online sources and, if necessary, patches them to make them more usable with Model Converter. @@ -129,11 +176,12 @@ above. -Convert the model to OpenVINO IR --------------------------------- +Convert the model to OpenVINO IR `⇑ <#top>`__ +############################################################################################################################### + ``omz_converter`` converts the models that are not in the OpenVINO™ IR -format into that format using Model Optimizer. +format into that format using model conversion API. The downloaded pytorch model is not in OpenVINO IR format which is required for inference with OpenVINO runtime. ``omz_converter`` is used @@ -155,40 +203,42 @@ respectively .. parsed-literal:: ========== Converting colorization-v2 to ONNX - Conversion to ONNX command: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/bin/python -- /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/model_zoo/internal_scripts/pytorch_to_onnx.py --model-path=models/public/colorization-v2 --model-name=ECCVGenerator --weights=models/public/colorization-v2/ckpt/colorization-v2-eccv16.pth --import-module=model --input-shape=1,1,256,256 --output-file=models/public/colorization-v2/colorization-v2-eccv16.onnx --input-names=data_l --output-names=color_ab + Conversion to ONNX command: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/bin/python -- /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/model_zoo/internal_scripts/pytorch_to_onnx.py --model-path=models/public/colorization-v2 --model-name=ECCVGenerator --weights=models/public/colorization-v2/ckpt/colorization-v2-eccv16.pth --import-module=model --input-shape=1,1,256,256 --output-file=models/public/colorization-v2/colorization-v2-eccv16.onnx --input-names=data_l --output-names=color_ab ONNX check passed successfully. ========== Converting colorization-v2 to IR (FP16) - Conversion command: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/bin/python -- /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/bin/mo --framework=onnx --output_dir=/tmp/tmp1_wcgtpa --model_name=colorization-v2 --input=data_l --output=color_ab --input_model=models/public/colorization-v2/colorization-v2-eccv16.onnx '--layout=data_l(NCHW)' '--input_shape=[1, 1, 256, 256]' --compress_to_fp16=True + Conversion command: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/bin/python -- /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/bin/mo --framework=onnx --output_dir=/tmp/tmp7wsuasz7 --model_name=colorization-v2 --input=data_l --output=color_ab --input_model=models/public/colorization-v2/colorization-v2-eccv16.onnx '--layout=data_l(NCHW)' '--input_shape=[1, 1, 256, 256]' --compress_to_fp16=True [ INFO ] Generated IR will be compressed to FP16. If you get lower accuracy, please consider disabling compression by removing argument --compress_to_fp16 or set it to false --compress_to_fp16=False. - Find more information about compression to FP16 at https://docs.openvino.ai/latest/openvino_docs_MO_DG_FP16_Compression.html + Find more information about compression to FP16 at https://docs.openvino.ai/2023.0/openvino_docs_MO_DG_FP16_Compression.html [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. - Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/latest/openvino_2_0_transition_guide.html + Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html [ SUCCESS ] Generated IR version 11 model. - [ SUCCESS ] XML file: /tmp/tmp1_wcgtpa/colorization-v2.xml - [ SUCCESS ] BIN file: /tmp/tmp1_wcgtpa/colorization-v2.bin + [ SUCCESS ] XML file: /tmp/tmp7wsuasz7/colorization-v2.xml + [ SUCCESS ] BIN file: /tmp/tmp7wsuasz7/colorization-v2.bin -Loading the Model ------------------ +Loading the Model `⇑ <#top>`__ +############################################################################################################################### -Load the model in OpenVINO Runtime with ``ie.read_model`` and compile it -for the specified device with ``ie.compile_model``. + Load the model in OpenVINO Runtime with +``ie.read_model`` and compile it for the specified device with +``ie.compile_model``. .. code:: ipython3 - ie = Core() - model = ie.read_model(model=MODEL_PATH) - compiled_model = ie.compile_model(model=model, device_name=DEVICE) + core = Core() + model = core.read_model(model=MODEL_PATH) + compiled_model = core.compile_model(model=model, device_name=device.value) input_layer = compiled_model.input(0) output_layer = compiled_model.output(0) N, C, H, W = list(input_layer.shape) -Utility Functions ------------------ +Utility Functions `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -264,8 +314,9 @@ Utility Functions plt.show() -Load the Image --------------- +Load the Image `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -330,8 +381,9 @@ Load the Image color_img_0 = colorize(test_img_0) color_img_1 = colorize(test_img_1) -Display Colorized Image ------------------------ +Display Colorized Image `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -339,7 +391,7 @@ Display Colorized Image -.. image:: 222-vision-image-colorization-with-output_files/222-vision-image-colorization-with-output_18_0.png +.. image:: 222-vision-image-colorization-with-output_files/222-vision-image-colorization-with-output_20_0.png .. code:: ipython3 @@ -348,5 +400,5 @@ Display Colorized Image -.. image:: 222-vision-image-colorization-with-output_files/222-vision-image-colorization-with-output_19_0.png +.. image:: 222-vision-image-colorization-with-output_files/222-vision-image-colorization-with-output_21_0.png diff --git a/docs/notebooks/222-vision-image-colorization-with-output_files/222-vision-image-colorization-with-output_18_0.png b/docs/notebooks/222-vision-image-colorization-with-output_files/222-vision-image-colorization-with-output_20_0.png similarity index 100% rename from docs/notebooks/222-vision-image-colorization-with-output_files/222-vision-image-colorization-with-output_18_0.png rename to docs/notebooks/222-vision-image-colorization-with-output_files/222-vision-image-colorization-with-output_20_0.png diff --git a/docs/notebooks/222-vision-image-colorization-with-output_files/222-vision-image-colorization-with-output_19_0.png b/docs/notebooks/222-vision-image-colorization-with-output_files/222-vision-image-colorization-with-output_21_0.png similarity index 100% rename from docs/notebooks/222-vision-image-colorization-with-output_files/222-vision-image-colorization-with-output_19_0.png rename to docs/notebooks/222-vision-image-colorization-with-output_files/222-vision-image-colorization-with-output_21_0.png diff --git a/docs/notebooks/222-vision-image-colorization-with-output_files/index.html b/docs/notebooks/222-vision-image-colorization-with-output_files/index.html index 9f0acbcf160..f7bb42accee 100644 --- a/docs/notebooks/222-vision-image-colorization-with-output_files/index.html +++ b/docs/notebooks/222-vision-image-colorization-with-output_files/index.html @@ -1,8 +1,8 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/222-vision-image-colorization-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/222-vision-image-colorization-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/222-vision-image-colorization-with-output_files/


../
-222-vision-image-colorization-with-output_18_0.png 12-Jul-2023 00:11              415792
-222-vision-image-colorization-with-output_19_0.png 12-Jul-2023 00:11              284966
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/222-vision-image-colorization-with-output_files/


../
+222-vision-image-colorization-with-output_20_0.png 16-Aug-2023 01:31              415792
+222-vision-image-colorization-with-output_21_0.png 16-Aug-2023 01:31              284966
 

diff --git a/docs/notebooks/223-text-prediction-with-output.rst b/docs/notebooks/223-text-prediction-with-output.rst index fed7cf99609..cb50b852e26 100644 --- a/docs/notebooks/223-text-prediction-with-output.rst +++ b/docs/notebooks/223-text-prediction-with-output.rst @@ -1,6 +1,8 @@ Text Prediction with OpenVINO™ ============================== +.. _top: + This notebook shows text prediction with OpenVINO. This notebook can work in two different modes, Text Generation and Conversation, which the user can select via selecting the model in the Model Selection Section. @@ -18,7 +20,7 @@ capabilities, including the ability to generate conditional synthetic text samples of unprecedented quality, where we prime the model with an input and have it generate a lengthy continuation. -More Details about the models are provided on their huggingface cards: +More details about the models are provided on their HuggingFace cards: - `GPT-2 `__ - `GPT-Neo `__ @@ -64,30 +66,60 @@ The following image illustrates the demo pipeline for conversation: image2 -For Conversation, User Input is tokenized with eos_token concatenated in -the end. Then, the text gets generated as detailed above. The Generated -response is added to the history with the eos_token at the end. -Additional user input is added to the history, and the sequence is -passed back into the model. +For Conversation, User Input is tokenized with ``eos_token`` +concatenated in the end. Then, the text gets generated as detailed +above. The Generated response is added to the history with the +``eos_token`` at the end. Additional user input is added to the history, +and the sequence is passed back into the model. + + +**Table of contents**: + +- `Model Selection <#model-selection>`__ +- `Load Model <#load-model>`__ + +- `Convert Pytorch Model to OpenVINO IR <#convert-pytorch-model-to-openvino-ir>`__ + + - `Load the model <#load-the-model>`__ + + - `Select inference device <#select-inference-device>`__ + +- `Pre-Processing <#pre-processing>`__ +- `Define tokenization <#define-tokenization>`__ + + - `Define Softmax layer <#define-softmax-layer>`__ + - `Set the minimum sequence length <#set-the-minimum-sequence-length>`__ + - `Top-K sampling <#top-k-sampling>`__ + - `Main Processing Function <#main-processing-function>`__ + +- `Inference with GPT-Neo/GPT-2 <#inference-with-gpt-neo-gpt-2>`__ +- `Conversation with PersonaGPT using OpenVINO™ <#conversation-with-personagpt-using-openvino>`__ +- `Converse Function <#converse-function>`__ +- `Conversation Class <#conversation-class>`__ +- `Conversation with PersonaGPT <#conversation-with-personagpt>`__ + +Model Selection `⇑ <#top>`__ +############################################################################################################################### -Model Selection ---------------- Select the Model to be used for text generation, GPT-2 and GPT-Neo are -used for text generation wheras PersonaGPT is used for Conversation. +used for text generation whereas PersonaGPT is used for Conversation. .. code:: ipython3 # Install Gradio for Interactive Inference and other requirements - !pip install -q 'openvino-dev>=2023.0.0' + !pip install -q "openvino-dev>=2023.0.0" !pip install -q gradio !pip install -q transformers[torch] onnx .. parsed-literal:: + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 ERROR: pip's dependency resolver does not currently take into account all the packages that are installed. This behaviour is the source of the following dependency conflicts. - pytorch-lightning 1.6.5 requires protobuf<=3.20.1, but you have protobuf 4.23.4 which is incompatible. + pytorch-lightning 1.6.5 requires protobuf<=3.20.1, but you have protobuf 4.24.0 which is incompatible. .. code:: ipython3 @@ -114,10 +146,10 @@ used for text generation wheras PersonaGPT is used for Conversation. -Load Model ----------- +Load Model `⇑ <#top>`__ +############################################################################################################################### -Download the Selected Model and Tokenizer from Huggingface +Download the Selected Model and Tokenizer from HuggingFace .. code:: ipython3 @@ -133,38 +165,40 @@ Download the Selected Model and Tokenizer from Huggingface pt_model = GPTNeoForCausalLM.from_pretrained('EleutherAI/gpt-neo-125M') tokenizer = GPT2TokenizerFast.from_pretrained('EleutherAI/gpt-neo-125M') -Convert Pytorch Model to OpenVINO IR ------------------------------------- +Convert Pytorch Model to OpenVINO IR `⇑ <#top>`__ +############################################################################################################################### + .. figure:: https://user-images.githubusercontent.com/29454499/211261803-784d4791-15cb-4aea-8795-0969dfbb8291.png :alt: conversion_pipeline conversion_pipeline -For starting work with GPT-Neo model using OpenVINO, model should be -converted to OpenVINO Intermediate Represenation (IR) format. -HuggingFace provides gpt-neo model in PyTorch format, which supported in -OpenVINO via conversion to ONNX. We use HuggingFace transformers -library’s oonx module to export model to ONNX. -``transformers.onnx.export`` accepts preprocessing function for input -sample generation (tokenizer in our case),an instance of model, ONNX -export configuration, ONNX opset version for export and output path. -More information about transformers export to ONNX can be found in -HuggingFace +For starting work with GPT-Neo model using OpenVINO, a model should be +converted to OpenVINO Intermediate Representation (IR) format. +HuggingFace provides a GPT-Neo model in PyTorch format, which is +supported in OpenVINO via conversion to ONNX. We use the HuggingFace +transformers library’s onnx module to export the model to ONNX. +``transformers.onnx.export`` accepts the preprocessing function for +input sample generation (the tokenizer in our case), an instance of the +model, ONNX export configuration, the ONNX opset version for export and +output path. More information about transformers export to ONNX can be +found in HuggingFace `documentation `__. While ONNX models are directly supported by OpenVINO runtime, it can be useful to convert them to IR format to take advantage of OpenVINO -optimization tools and features. ``mo.convert_model`` python function -can be used for converting model using `OpenVINO Model -Optimizer `__. -The function returns instance of OpenVINO Model class, which is ready to -use in Python interface but can also be serialized to OpenVINO IR format -for future execution using ``openvino.runtime.serialize``. In our case, -``compress_to_fp16`` parameter is enabled for compression model weights -to fp16 precision and also specified dynamic input shapes with possible -shape range (from 1 token to maximum length defined in our processing -function) for optimization of memory consumption. +optimization tools and features. The ``mo.convert_model`` Python +function of `model conversion +API `__ +can be used for converting the model. The function returns instance of +OpenVINO Model class, which is ready to use in Python interface but can +also be serialized to OpenVINO IR format for future execution using +``openvino.runtime.serialize``. In our case, the ``compress_to_fp16`` +parameter is enabled for compression model weights to FP16 precision and +also specified dynamic input shapes with a possible shape range (from 1 +token to a maximum length defined in our processing function) for +optimization of memory consumption. .. code:: ipython3 @@ -192,7 +226,7 @@ function) for optimization of memory consumption. # convert model to openvino if model_name.value == "PersonaGPT (Converastional)": - ov_model = mo.convert_model(onnx_path, compress_to_fp16=True, input="input_ids[1,1..1000],attention_mask[1,1..1000]") + ov_model = mo.convert_model(onnx_path, compress_to_fp16=True, input="input_ids[1,-1],attention_mask[1,-1]") else: ov_model = mo.convert_model(onnx_path, compress_to_fp16=True, input="input_ids[1,1..128],attention_mask[1,1..128]") @@ -202,38 +236,61 @@ function) for optimization of memory consumption. .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/gpt2/modeling_gpt2.py:810: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/gpt2/modeling_gpt2.py:807: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if batch_size <= 0: -Load the model -~~~~~~~~~~~~~~ +Load the model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + We start by building an OpenVINO Core object. Then we read the network -architecture and model weights from the .xml and .bin files, +architecture and model weights from the ``.xml`` and ``.bin`` files, respectively. Finally, we compile the model for the desired device. -Because we use the dynamic shapes feature, which is only available on -CPU, we must use ``CPU`` for the device. Dynamic shapes support on GPU -is coming soon. -Since the text recognition model has a dynamic input shape, you cannot -directly switch device to ``GPU`` for inference on integrated or -discrete Intel GPUs. In order to run inference on iGPU or dGPU with this -model, you will need to resize the inputs to this model to use a fixed -size and then try running the inference on ``GPU`` device. +Select inference device `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + + +Select device from dropdown list for running inference using OpenVINO: .. code:: ipython3 from openvino.runtime import Core + import ipywidgets as widgets + core = Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + # initialize openvino core core = Core() # read the model and corresponding weights from file model = core.read_model(model_path) - + +.. code:: ipython3 + # compile the model for CPU devices - compiled_model = core.compile_model(model=model, device_name="CPU") + compiled_model = core.compile_model(model=model, device_name=device.value) # get output tensors output_key = compiled_model.output(0) @@ -243,16 +300,18 @@ names of the output nodes of the network. In the case of GPT-Neo, we have ``batch size`` and ``sequence length`` as inputs and ``batch size``, ``sequence length`` and ``vocab size`` as outputs. -Pre-Processing --------------- +Pre-Processing `⇑ <#top>`__ +############################################################################################################################### + NLP models often take a list of tokens as a standard input. A token is a word or a part of a word mapped to an integer. To provide the proper input, we use a vocabulary file to handle the mapping. So first let’s load the vocabulary file. -Define tokenization -------------------- +Define tokenization `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -283,11 +342,10 @@ at later stage. eos_token_id = tokenizer.eos_token_id eos_token = tokenizer.decode(eos_token_id) -Define Softmax layer -~~~~~~~~~~~~~~~~~~~~ +Define Softmax layer `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -A softmax function is used to convert top-k logits into a probability -distribution. +A softmax function is used to convert top-k logits into a probability distribution. .. code:: ipython3 @@ -299,12 +357,12 @@ distribution. summation = e_x.sum(axis=-1, keepdims=True) return e_x / summation -Set the minimum sequence length -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Set the minimum sequence length `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -If the minimum sequence length is not reached, the following code will -reduce the probability of the ``eos`` token occurring. This continues -the process of generating the next words. +If the minimum sequence length is not reached, the following code will reduce the probability of +the ``eos`` token occurring. This continues the process of generating +the next words. .. code:: ipython3 @@ -325,11 +383,11 @@ the process of generating the next words. scores[:, eos_token_id] = -float("inf") return scores -Top-K sampling -~~~~~~~~~~~~~~ +Top-K sampling `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -In Top-K sampling, we filter the K most likely next words and -redistribute the probability mass among only those K next words. +In Top-K sampling, we filter the K most likely next words and redistribute the probability mass among only those +K next words. .. code:: ipython3 @@ -354,8 +412,8 @@ redistribute the probability mass among only those K next words. fill_value=filter_value).filled() return filtred_scores -Main Processing Function -~~~~~~~~~~~~~~~~~~~~~~~~ +Main Processing Function `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ Generating the predicted sequence. @@ -404,11 +462,10 @@ Generating the predicted sequence. attention_mask = np.concatenate((attention_mask, [[1] * len(next_tokens)]), axis=-1) return input_ids -Inference with GPT-Neo/GPT-2 ----------------------------- +Inference with GPT-Neo/GPT-2 `⇑ <#top>`__ +############################################################################################################################### -The ``text`` variable below is the input used to generate a predicted -sequence. +The ``text`` variable below is the input used to generate a predicted sequence. .. code:: ipython3 @@ -437,23 +494,23 @@ sequence. Selected Model is PersonaGPT. Please select GPT-Neo or GPT-2 in the first cell to generate text sequences -Conversation with PersonaGPT using OpenVINO™ -============================================ +# Conversation with PersonaGPT using OpenVINO™ `⇑ <#top>`__ -User Input is tokenized with eos_token concatenated in the end. Model -input is tokenized text, which serves as initial condition for +User Input is tokenized with ``eos_token`` concatenated in the end. +Model input is tokenized text, which serves as initial condition for generation, then logits from model inference result should be obtained and token with the highest probability is selected using top-k sampling strategy and joined to input sequence. The procedure repeats until end -of sequence token will be recived or specified maximum length is +of sequence token will be received or specified maximum length is reached. After that, decoding token ids to text using tokenized should be applied. -The Generated response is added to the history with the eos_token at the -end. Further User Input is added to it and agin passed into the model. +The Generated response is added to the history with the ``eos_token`` at +the end. Further User Input is added to it and again passed into the +model. -Converse Function ------------------ +Converse Function `⇑ <#top>`__ +############################################################################################################################### Wrapper on generate sequence function to support conversation @@ -497,8 +554,9 @@ Wrapper on generate sequence function to support conversation response = ''.join(tokenizer.batch_decode(history)).split(eos_token)[-2] return response, history -Conversation Class ------------------- +Conversation Class `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -521,8 +579,9 @@ Conversation Class self.messages.append(f"PersonaGPT: {response}") return response -Conversation with PersonaGPT ----------------------------- +Conversation with PersonaGPT `⇑ <#top>`__ +############################################################################################################################### + This notebook provides two styles of inference, Plain and Interactive. The style of inference can be selected in the next cell. @@ -601,23 +660,23 @@ The style of inference can be selected in the next cell. .. parsed-literal:: Person: Hi,How are you? - PersonaGPT: good, how are you doing? + PersonaGPT: good, how about you? what do you like to do for fun? Person: What are you doing? - PersonaGPT: i'm good thanks what are you up too + PersonaGPT: i'm playing some video games. Person: I like to dance,do you? - PersonaGPT: i like to read books + PersonaGPT: i don't have any dancing abilities. Person: Can you recommend me some books? - PersonaGPT: yes i can i like books about dance + PersonaGPT: anybody can do it if you try. Person: Hi,How are you? - PersonaGPT: i am good thanks for asking + PersonaGPT: good, do you have any hobbies? Person: What are you doing? - PersonaGPT: i'm just sitting at home reading + PersonaGPT: i love to cook. Person: I like to dance,do you? - PersonaGPT: no but i love reading + PersonaGPT: i don't have any musical abilities. Person: Can you recommend me some books? - PersonaGPT: yes i like to read too + PersonaGPT: anybody can do it if you try. Person: Hi,How are you? - PersonaGPT: good. do you like to cook? + PersonaGPT: good, do you like cooking? Person: What are you doing? - PersonaGPT: i'm cooking right now. + PersonaGPT: i am watching netflix. diff --git a/docs/notebooks/224-3D-segmentation-point-clouds-with-output.rst b/docs/notebooks/224-3D-segmentation-point-clouds-with-output.rst index 42311dc51ed..fef333d4d1c 100644 --- a/docs/notebooks/224-3D-segmentation-point-clouds-with-output.rst +++ b/docs/notebooks/224-3D-segmentation-point-clouds-with-output.rst @@ -1,6 +1,8 @@ Part Segmentation of 3D Point Clouds with OpenVINO™ =================================================== +.. _top: + This notebook demonstrates how to process `point cloud `__ data and run 3D Part Segmentation with OpenVINO. We use the @@ -8,12 +10,12 @@ Part Segmentation with OpenVINO. We use the detect each part of a chair and return its category. PointNet -######## +-------- PointNet was proposed by Charles Ruizhongtai Qi, a researcher at -Stanford University in 2016: arXiv:1612.00593 <`PointNet: Deep Learning -on Point Sets for 3D Classification and -Segmentation `__>. The motivation +Stanford University in 2016: `PointNet: Deep Learning on Point Sets for +3D Classification and +Segmentation `__. The motivation behind the research is to classify and segment 3D representations of images. They use a data structure called point cloud, which is a set of points that represents a 3D shape or object. PointNet provides a unified @@ -22,8 +24,19 @@ segmentation, to scene semantic parsing. It is highly efficient and effective, showing strong performance on par or even better than state of the art. -Imports -------- +**Table of contents**: + +- `Imports <#imports>`__ +- `Prepare the Model <#prepare-the-model>`__ +- `Data Processing Module <#data-processing-module>`__ +- `Visualize the original 3D data <#visualize-the-original-3d-data>`__ +- `Run inference <#run-inference>`__ + + - `Select inference device <#select-inference-device>`__ + +Imports `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -39,12 +52,12 @@ Imports sys.path.append("../utils") from notebook_utils import download_file -Prepare the Model ------------------ +Prepare the Model `⇑ <#top>`__ +############################################################################################################################### -Download the pre-trained PointNet ONNX model. This pre-trained model is -provided by `axinc-ai `__, and you can find -more point clouds examples +Download the pre-trained PointNet ONNX model. This pre-trained model is provided by +`axinc-ai `__, and you can find more +point clouds examples `here `__. .. code:: ipython3 @@ -58,19 +71,19 @@ more point clouds examples Convert the ONNX model to OpenVINO IR. An OpenVINO IR (Intermediate Representation) model consists of an ``.xml`` file, containing information about network topology, and a ``.bin`` file, containing the -weights and biases binary data. Model Optimizer Python API used for -conversion ONNX model to OpenVINO IR. The ``mo.convert_model`` Python -function returns an OpenVINO model ready to load on device and start -making predictions. We can save it on disk for next usage with -``openvino.runtime.serialize``. For more information about Model -Optimizer Python API, see the `Model Optimizer Developer -Guide `__. +weights and biases binary data. Model conversion Python API is used for +conversion of ONNX model to OpenVINO IR. The ``mo.convert_model`` Python +function returns an OpenVINO model ready to load on a device and start +making predictions. We can save it on a disk for next usage with +``openvino.runtime.serialize``. For more information about model +conversion Python API, see this +`page `__. .. code:: ipython3 ir_model_xml = onnx_model_path.with_suffix(".xml") - ie = Core() + core = Core() if not ir_model_xml.exists(): # Convert model to OpenVINO Model @@ -79,11 +92,12 @@ Guide `__. serialize(model, str(ir_model_xml)) else: # Read model - model = ie.read_model(model=ir_model_xml) + model = core.read_model(model=ir_model_xml) -Data Processing Module ----------------------- +Data Processing Module `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -137,8 +151,8 @@ Data Processing Module return ax -Visualize the original 3D data ------------------------------- +Visualize the original 3D data `⇑ <#top>`__ +############################################################################################################################### The point cloud data can be downloaded from `ShapeNet `__, @@ -163,14 +177,14 @@ chair for example. .. image:: 224-3D-segmentation-point-clouds-with-output_files/224-3D-segmentation-point-clouds-with-output_10_0.png -Run inference -------------- +Run inference `⇑ <#top>`__ +############################################################################################################################### -Run inference and visualize the results of 3D segmentation. - The input -data is a point cloud with ``1 batch size``\ ,\ ``3 axis value`` (x, y, -z) and ``arbitrary number of points`` (dynamic shape). - The output data -is a mask with ``1 batch size`` and ``4 classification confidence`` for -each input point. +Run inference and visualize the results of 3D segmentation. - The input data is a point cloud with +``1 batch size``\ ,\ ``3 axis value`` (x, y, z) and +``arbitrary number of points`` (dynamic shape). - The output data is a +mask with ``1 batch size`` and ``4 classification confidence`` for each +input point. .. code:: ipython3 @@ -192,10 +206,38 @@ each input point. output shape: [1,?,4] +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + .. code:: ipython3 # Inference - compiled_model = ie.compile_model(model=model, device_name="CPU") + compiled_model = core.compile_model(model=model, device_name=device.value) output_layer = compiled_model.output(0) result = compiled_model([point])[output_layer] @@ -224,5 +266,5 @@ each input point. -.. image:: 224-3D-segmentation-point-clouds-with-output_files/224-3D-segmentation-point-clouds-with-output_13_0.png +.. image:: 224-3D-segmentation-point-clouds-with-output_files/224-3D-segmentation-point-clouds-with-output_15_0.png diff --git a/docs/notebooks/224-3D-segmentation-point-clouds-with-output_files/224-3D-segmentation-point-clouds-with-output_13_0.png b/docs/notebooks/224-3D-segmentation-point-clouds-with-output_files/224-3D-segmentation-point-clouds-with-output_15_0.png similarity index 100% rename from docs/notebooks/224-3D-segmentation-point-clouds-with-output_files/224-3D-segmentation-point-clouds-with-output_13_0.png rename to docs/notebooks/224-3D-segmentation-point-clouds-with-output_files/224-3D-segmentation-point-clouds-with-output_15_0.png diff --git a/docs/notebooks/224-3D-segmentation-point-clouds-with-output_files/index.html b/docs/notebooks/224-3D-segmentation-point-clouds-with-output_files/index.html index ae1286a34a7..53f958fc341 100644 --- a/docs/notebooks/224-3D-segmentation-point-clouds-with-output_files/index.html +++ b/docs/notebooks/224-3D-segmentation-point-clouds-with-output_files/index.html @@ -1,8 +1,8 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/224-3D-segmentation-point-clouds-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/224-3D-segmentation-point-clouds-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/224-3D-segmentation-point-clouds-with-output_files/


../
-224-3D-segmentation-point-clouds-with-output_10..> 12-Jul-2023 00:11              209355
-224-3D-segmentation-point-clouds-with-output_13..> 12-Jul-2023 00:11              222552
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/224-3D-segmentation-point-clouds-with-output_files/


../
+224-3D-segmentation-point-clouds-with-output_10..> 16-Aug-2023 01:31              209355
+224-3D-segmentation-point-clouds-with-output_15..> 16-Aug-2023 01:31              222552
 

diff --git a/docs/notebooks/225-stable-diffusion-text-to-image-with-output.rst b/docs/notebooks/225-stable-diffusion-text-to-image-with-output.rst index 1802a6d48bb..812d448e31d 100644 --- a/docs/notebooks/225-stable-diffusion-text-to-image-with-output.rst +++ b/docs/notebooks/225-stable-diffusion-text-to-image-with-output.rst @@ -1,6 +1,8 @@ Text-to-Image Generation with Stable Diffusion and OpenVINO™ ============================================================ +.. _top: + Stable Diffusion is a text-to-image latent diffusion model created by the researchers and engineers from `CompVis `__, `Stability @@ -35,11 +37,29 @@ using OpenVINO. Notebook contains the following steps: 1. Convert PyTorch models to ONNX format. -2. Convert ONNX models to OpenVINO IR format, using Model Optimizer tool. +2. Convert ONNX models to OpenVINO IR format, using model conversion + API. 3. Run Stable Diffusion pipeline with OpenVINO. -Prerequisites -------------- +**Table of contents**: + +- `Prerequisites <#prerequisites>`__ +- `Create PyTorch Models pipeline <#create-pytorch-models-pipeline>`__ +- `Convert models to OpenVINO Intermediate representation (IR) format <#convert-models-to-openvino-intermediate-representation-ir-format>`__ + + - `Text Encoder <#text-encoder>`__ + - `U-net <#u-net>`__ + - `VAE <#vae>`__ + +- `Prepare Inference Pipeline <#prepare-inference-pipeline>`__ +- `Configure Inference Pipeline <#configure-inference-pipeline>`__ + + - `Text-to-Image generation <#text-to-image-generation>`__ + - `Image-to-Image generation <#image-to-image-generation>`__ + +Prerequisites `⇑ <#top>`__ +############################################################################################################################### + **The following is needed only if you want to use the original model. If not, you do not have to do anything. Just run the notebook.** @@ -58,7 +78,8 @@ not, you do not have to do anything. Just run the notebook.** .. code:: python - ## login to huggingfacehub to get access to pretrained model + + ## login to huggingfacehub to get access to pretrained model from huggingface_hub import notebook_login, whoami try: @@ -76,15 +97,15 @@ solutions based on Stable Diffusion. .. code:: ipython3 - !pip install -q 'diffusers[torch]>=0.9.0' - !pip install -q 'huggingface-hub>=0.9.1' + !pip install -q "diffusers[torch]>=0.9.0" + !pip install -q "huggingface-hub>=0.9.1" -Create Pytorch Models pipeline ------------------------------- +Create PyTorch Models pipeline `⇑ <#top>`__ +############################################################################################################################### -StableDiffusionPipeline is an end-to-end inference pipeline that you can -use to generate images from text with just a few lines of code. +``StableDiffusionPipeline`` is an end-to-end inference pipeline that you can use to generate images +from text with just a few lines of code. First, load the pre-trained weights of all components of the model. @@ -109,8 +130,8 @@ First, load the pre-trained weights of all components of the model. Fetching 15 files: 0%| | 0/15 [00:00`__ +############################################################################################################################### OpenVINO supports PyTorch through export to the ONNX format. You will use ``torch.onnx.export`` function for obtaining ONNX model. You can @@ -123,21 +144,23 @@ input and output names or dynamic shapes). While ONNX models are directly supported by OpenVINO™ runtime, it can be useful to convert them to IR format to take advantage of advanced -OpenVINO optimization tools and features. You will use OpenVINO Model -Optimizer tool for conversion model to IR format and compression weights -to ``FP16`` format. +OpenVINO optimization tools and features. For converting the model to IR +format and compressing weights to ``FP16`` format, you will use model +conversion API. The model consists of three important parts: -* Text Encoder for creation condition to generate image from text prompt. -* Unet for step by step denoising latent image representation. -* Autoencoder (VAE) for encdoing input image to latent space (if required) and decoding latent -space to image back after generation. +- Text Encoder for creation condition to generate image from text + prompt. +- Unet for step by step denoising latent image representation. +- Autoencoder (VAE) for encoding input image to latent space (if + required) and decoding latent space to image back after generation. Let us convert each part. -Text Encoder -~~~~~~~~~~~~ +Text Encoder `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The text-encoder is responsible for transforming the input prompt, for example, “a photo of an astronaut riding a horse” into an embedding @@ -218,16 +241,16 @@ hidden states. You will use ``opset_version=14``, because model contains -U-net -~~~~~ +U-net `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Unet model has three inputs: -* ``sample`` - latent image sample from -previous step. Generation process has not been started yet, so you will -use random noise. -* ``timestep`` - current scheduler step. -* ``encoder_hidden_state`` - hidden state of text encoder. +- ``sample`` - latent image sample from previous step. Generation + process has not been started yet, so you will use random noise. +- ``timestep`` - current scheduler step. +- ``encoder_hidden_state`` - hidden state of text encoder. Model predicts the ``sample`` state for the next step. @@ -295,8 +318,9 @@ Model predicts the ``sample`` state for the next step. -VAE -~~~ +VAE `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The VAE model has two parts, an encoder and a decoder. The encoder is used to convert the image into a low dimensional latent representation, @@ -410,13 +434,17 @@ of the pipeline, it will be better to convert them to separate models. VAE decoder will be loaded from vae_decoder.xml -Prepare Inference Pipeline --------------------------- +Prepare Inference Pipeline `⇑ <#top>`__ +############################################################################################################################### + Putting it all together, let us now take a closer look at how the model works in inference by illustrating the logical flow. -.. image:: https://camo.githubusercontent.com/60d9edf7fc65d617c56ddac5344a9fe3f4152e38f43cab994589a8e8bf14d36a/68747470733a2f2f757365722d696d616765732e67697468756275736572636f6e74656e742e636f6d2f32393435343439392f3231363337383933322d37613962653339662d636338362d343365342d623037322d3636333732613335643662642e706e67 +.. figure:: https://user-images.githubusercontent.com/29454499/216378932-7a9be39f-cc86-43e4-b072-66372a35d6bd.png + :alt: sd-pipeline + + sd-pipeline As you can see from the diagram, the only difference between Text-to-Image and text-guided Image-to-Image generation in approach is @@ -770,8 +798,9 @@ of the variational auto encoder. return timesteps, num_inference_steps - t_start -Configure Inference Pipeline ----------------------------- +Configure Inference Pipeline `⇑ <#top>`__ +############################################################################################################################### + First, you should create instances of OpenVINO Model. @@ -779,16 +808,35 @@ First, you should create instances of OpenVINO Model. from openvino.runtime import Core core = Core() - text_enc = core.compile_model(TEXT_ENCODER_OV_PATH, 'AUTO') + +Select device from dropdown list for running inference using OpenVINO. .. code:: ipython3 - unet_model = core.compile_model(UNET_OV_PATH, 'AUTO') + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device .. code:: ipython3 - vae_decoder = core.compile_model(VAE_DECODER_OV_PATH, 'AUTO') - vae_encoder = core.compile_model(VAE_ENCODER_OV_PATH, 'AUTO') + + text_enc = core.compile_model(TEXT_ENCODER_OV_PATH, device.value) + +.. code:: ipython3 + + unet_model = core.compile_model(UNET_OV_PATH, device.value) + +.. code:: ipython3 + + vae_decoder = core.compile_model(VAE_DECODER_OV_PATH, device.value) + vae_encoder = core.compile_model(VAE_ENCODER_OV_PATH, device.value) Model tokenizer and scheduler are also important parts of the pipeline. Let us define them and put all components together @@ -814,16 +862,16 @@ Let us define them and put all components together scheduler=lms ) -Text-to-Image generation -~~~~~~~~~~~~~~~~~~~~~~~~ +Text-to-Image generation `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Now, you can define a text prompt for image generation and run inference pipeline. Optionally, you can also change the random generator seed for -latent state initialization and number of steps. +latent state initialization and number of steps. -.. note:: - - Consider increasing ``steps`` to get more precise results. A suggested value is ``50``, but it will take longer time to process. + **Note**: Consider increasing ``steps`` to get more precise results. + A suggested value is ``50``, but it will take longer time to process. .. code:: ipython3 @@ -907,13 +955,14 @@ Now is show time! -.. image:: 225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_29_1.png +.. image:: 225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_33_1.png Nice. As you can see, the picture has quite a high definition 🔥. -Image-to-Image generation -~~~~~~~~~~~~~~~~~~~~~~~~~ +Image-to-Image generation `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Image-to-Image generation, additionally to text prompt, requires providing initial image. Optionally, you can also change ``strength`` @@ -972,7 +1021,7 @@ semantically consistent with the input. -.. image:: 225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_33_1.png +.. image:: 225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_37_1.png @@ -1005,5 +1054,5 @@ semantically consistent with the input. -.. image:: 225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_35_1.png +.. image:: 225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_39_1.png diff --git a/docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_29_1.png b/docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_29_1.png deleted file mode 100644 index 1052cb35d05..00000000000 --- a/docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_29_1.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:d0f1bbca67c22a713aad717ea3ee14fb453e85d45d542429059bb25fee2b7d6d -size 372493 diff --git a/docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_33_1.png b/docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_33_1.png index 431d6ff733b..1052cb35d05 100644 --- a/docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_33_1.png +++ b/docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_33_1.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:d8a9811dc8ab8b5f2ef6171ba757ccf9bbe33d0f137824ad0c32380bd74c5ff4 -size 928896 +oid sha256:d0f1bbca67c22a713aad717ea3ee14fb453e85d45d542429059bb25fee2b7d6d +size 372493 diff --git a/docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_37_1.png b/docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_37_1.png new file mode 100644 index 00000000000..431d6ff733b --- /dev/null +++ b/docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_37_1.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:d8a9811dc8ab8b5f2ef6171ba757ccf9bbe33d0f137824ad0c32380bd74c5ff4 +size 928896 diff --git a/docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_35_1.png b/docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_39_1.png similarity index 100% rename from docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_35_1.png rename to docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/225-stable-diffusion-text-to-image-with-output_39_1.png diff --git a/docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/index.html b/docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/index.html index f9dae8a6c49..c148f018de1 100644 --- a/docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/index.html +++ b/docs/notebooks/225-stable-diffusion-text-to-image-with-output_files/index.html @@ -1,9 +1,9 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/225-stable-diffusion-text-to-image-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/225-stable-diffusion-text-to-image-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/225-stable-diffusion-text-to-image-with-output_files/


../
-225-stable-diffusion-text-to-image-with-output_..> 12-Jul-2023 00:11              372493
-225-stable-diffusion-text-to-image-with-output_..> 12-Jul-2023 00:11              928896
-225-stable-diffusion-text-to-image-with-output_..> 12-Jul-2023 00:11              726937
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/225-stable-diffusion-text-to-image-with-output_files/


../
+225-stable-diffusion-text-to-image-with-output_..> 16-Aug-2023 01:31              372493
+225-stable-diffusion-text-to-image-with-output_..> 16-Aug-2023 01:31              928896
+225-stable-diffusion-text-to-image-with-output_..> 16-Aug-2023 01:31              726937
 

diff --git a/docs/notebooks/226-yolov7-optimization-with-output.rst b/docs/notebooks/226-yolov7-optimization-with-output.rst index 13440809f51..c04fb0c6263 100644 --- a/docs/notebooks/226-yolov7-optimization-with-output.rst +++ b/docs/notebooks/226-yolov7-optimization-with-output.rst @@ -1,6 +1,8 @@ Convert and Optimize YOLOv7 with OpenVINO™ ========================================== +.. _top: + The YOLOv7 algorithm is making big waves in the computer vision and machine learning communities. It is a real-time object detection algorithm that performs image recognition tasks by taking an image as @@ -24,37 +26,66 @@ include video analytics, robotics, autonomous vehicles, multi-object tracking and object counting, medical image analysis, and many others. This tutorial demonstrates step-by-step instructions on how to run and -optimize PyTorch Yolo V7 with OpenVINO. +optimize PyTorch YOLO V7 with OpenVINO. The tutorial consists of the following steps: -- Prepare PyTorch model -- Download and prepare dataset -- Validate original model -- Convert PyTorch model to ONNX -- Convert ONNX model to OpenVINO IR -- Validate converted model -- Prepare and run optimization pipeline -- Compare accuracy of the FP32 and quantized models. -- Compare performance of the FP32 and quantized models. +- Prepare PyTorch model +- Download and prepare dataset +- Validate original model +- Convert PyTorch model to ONNX +- Convert ONNX model to OpenVINO IR +- Validate converted model +- Prepare and run optimization pipeline +- Compare accuracy of the FP32 and quantized models. +- Compare performance of the FP32 and quantized models. + +**Table of contents**: + +- `Get Pytorch model <#get-pytorch-model>`__ +- `Prerequisites <#prerequisites>`__ +- `Check model inference <#check-model-inference>`__ +- `Export to ONNX <#export-to-onnx>`__ +- `Convert ONNX Model to OpenVINO Intermediate Representation (IR) <#convert-onnx-model-to-openvino-intermediate-representation-ir>`__ +- `Verify model inference <#verify-model-inference>`__ + + - `Preprocessing <#preprocessing>`__ + - `Postprocessing <#postprocessing>`__ + - `Select inference device <#select-inference-device>`__ + +- `Verify model accuracy <#verify-model-accuracy>`__ + + - `Download dataset <#download-dataset>`__ + - `Create dataloader <#create-dataloader>`__ + - `Define validation function <#define-validation-function>`__ + +- `Optimize model using NNCF Post-training Quantization API <#optimize-model-using-nncf-post-training-quantization-api>`__ +- `Validate Quantized model inference <#validate-quantized-model-inference>`__ +- `Validate quantized model accuracy <#validate-quantized-model-accuracy>`__ +- `Compare Performance of the Original and Quantized Models <#compare-performance-of-the-original-and-quantized-models>`__ + +Get Pytorch model `⇑ <#top>`__ +############################################################################################################################### -Get Pytorch model ------------------ Generally, PyTorch models represent an instance of the `torch.nn.Module `__ class, initialized by a state dictionary with model weights. We will use the YOLOv7 tiny model pre-trained on a COCO dataset, which is available in this `repo `__. Typical steps -to obtain pre-trained model: 1. Create instance of model class. 2. Load -checkpoint state dict, which contains pre-trained model weights. 3. Turn -model to evaluation for switching some operations to inference mode. +to obtain pre-trained model: + +1. Create instance of model class. +2. Load checkpoint state dict, which contains pre-trained model weights. +3. Turn model to evaluation for switching some operations to inference + mode. In this case, the model creators provide a tool that enables converting the YOLOv7 model to ONNX, so we do not need to do these steps manually. -Prerequisites -------------- +Prerequisites `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -78,9 +109,9 @@ Prerequisites remote: Counting objects: 100% (6/6), done. remote: Compressing objects: 100% (4/4), done. remote: Total 1191 (delta 2), reused 6 (delta 2), pack-reused 1185 - Receiving objects: 100% (1191/1191), 74.23 MiB | 4.24 MiB/s, done. + Receiving objects: 100% (1191/1191), 74.23 MiB | 4.20 MiB/s, done. Resolving deltas: 100% (511/511), done. - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/226-yolov7-optimization/yolov7 + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/226-yolov7-optimization/yolov7 .. code:: ipython3 @@ -105,12 +136,13 @@ Prerequisites .. parsed-literal:: - PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/226-yolov7-optimization/yolov7/model/yolov7-tiny.pt') + PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/226-yolov7-optimization/yolov7/model/yolov7-tiny.pt') -Check model inference ---------------------- +Check model inference `⇑ <#top>`__ +############################################################################################################################### + ``detect.py`` script run pytorch model inference and save image as result, @@ -131,7 +163,7 @@ result, traced_script_module saved! model is traced! - 5 horses, Done. (70.2ms) Inference, (0.8ms) NMS + 5 horses, Done. (70.8ms) Inference, (0.8ms) NMS The image with the result is saved in: runs/detect/exp/horses.jpg Done. (0.084s) @@ -145,12 +177,13 @@ result, -.. image:: 226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_8_0.png +.. image:: 226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_9_0.png -Export to ONNX --------------- +Export to ONNX `⇑ <#top>`__ +############################################################################################################################### + To export an ONNX format of the model, we will use ``export.py`` script. Let us check its arguments. @@ -196,19 +229,19 @@ Let us check its arguments. The most important parameters: -* ``--weights`` - path to model weigths checkpoint -* ``--img-size`` - size of input image for onnx tracing +- ``--weights`` - path to model weights checkpoint +- ``--img-size`` - size of input image for onnx tracing When exporting the ONNX model from PyTorch, there is an opportunity to setup configurable parameters for including post-processing results in model: -* ``--end2end`` - export full model to onnx including post-processing -* ``--grid`` - export Detect layer as part of model -* ``--topk-all`` - topk elements for all images -* ``--iou-thres`` - intersection over union threshold for NMS -* ``--conf-thres`` - minimal confidence threshold -* ``--max-wh`` - max bounding box width and height for NMS +- ``--end2end`` - export full model to onnx including post-processing +- ``--grid`` - export Detect layer as part of model +- ``--topk-all`` - top k elements for all images +- ``--iou-thres`` - intersection over union threshold for NMS +- ``--conf-thres`` - minimal confidence threshold +- ``--max-wh`` - max bounding box width and height for NMS Including whole post-processing to model can help to achieve more performant results, but in the same time it makes the model less @@ -242,19 +275,19 @@ an end2end ONNX model, you can check this Starting ONNX export with onnx 1.14.0... ONNX export success, saved as model/yolov7-tiny.onnx - Export complete (2.52s). Visualize with https://github.com/lutzroeder/netron. + Export complete (2.53s). Visualize with https://github.com/lutzroeder/netron. -Convert ONNX Model to OpenVINO Intermediate Representation (IR) ---------------------------------------------------------------- +Convert ONNX Model to OpenVINO Intermediate Representation (IR). `⇑ <#top>`__ +############################################################################################################################### -While ONNX models are directly supported by OpenVINO runtime, it can be -useful to convert them to IR format to take the advantage of OpenVINO -optimization tools and features. The ``mo.convert_model`` python -function in OpenVINO Model Optimizer can be used for converting the -model. The function returns instance of OpenVINO Model class, which is -ready to use in Python interface. However, it can also be serialized to -OpenVINO IR format for future execution. +While ONNX models are directly supported by OpenVINO runtime, +it can be useful to convert them to IR format to take the advantage of +OpenVINO optimization tools and features. The ``mo.convert_model`` +python function in OpenVINO Model Optimizer can be used for converting +the model. The function returns instance of OpenVINO Model class, which +is ready to use in Python interface. However, it can also be serialized +to OpenVINO IR format for future execution. .. code:: ipython3 @@ -265,23 +298,25 @@ OpenVINO IR format for future execution. # serialize model for saving IR serialize(model, 'model/yolov7-tiny.xml') -Verify model inference ----------------------- +Verify model inference `⇑ <#top>`__ +############################################################################################################################### + To test model work, we create inference pipeline similar to ``detect.py``. The pipeline consists of preprocessing step, inference of OpenVINO model, and results post-processing to get bounding boxes. -Preprocessing -~~~~~~~~~~~~~ +Preprocessing `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Model input is a tensor with the ``[1, 3, 640, 640]`` shape in ``N, C, H, W`` format, where -* ``N`` - number of images in batch (batch size) -* ``C`` - image channels -* ``H`` - image height -* ``W`` - image width +- ``N`` - number of images in batch (batch size) +- ``C`` - image channels +- ``H`` - image height +- ``W`` - image width Model expects images in RGB channels format and normalized in [0, 1] range. To resize images to fit model size ``letterbox`` resize approach @@ -352,8 +387,9 @@ To keep specific shape, preprocessing automatically enables padding. COLORS = {name: [np.random.randint(0, 255) for _ in range(3)] for i, name in enumerate(NAMES)} -Postprocessing -~~~~~~~~~~~~~~ +Postprocessing `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Model output contains detection boxes candidates. It is a tensor with the ``[1,25200,85]`` shape in the ``B, N, 85`` format, where: @@ -432,8 +468,39 @@ algorithm and rescale boxes coordinates to original image size. core = Core() # read converted model model = core.read_model('model/yolov7-tiny.xml') + +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + # load model on CPU device - compiled_model = core.compile_model(model, 'CPU') + compiled_model = core.compile_model(model, device.value) .. code:: ipython3 @@ -445,15 +512,17 @@ algorithm and rescale boxes coordinates to original image size. -.. image:: 226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_22_0.png +.. image:: 226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_26_0.png -Verify model accuracy ---------------------- +Verify model accuracy `⇑ <#top>`__ +############################################################################################################################### + + +Download dataset `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Download dataset -~~~~~~~~~~~~~~~~ YOLOv7 tiny is pre-trained on the COCO dataset, so in order to evaluate the model accuracy, we need to download it. According to the @@ -495,8 +564,9 @@ the original model evaluation scripts. coco2017labels-segments.zip: 0%| | 0.00/169M [00:00`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -522,11 +592,12 @@ Create dataloader .. parsed-literal:: - val: Scanning 'coco/val2017' images and labels... 4952 found, 48 missing, 0 empty, 0 corrupted: 100%|██████████| 5000/5000 [00:01<00:00, 2847.41it/s] + val: Scanning 'coco/val2017' images and labels... 4952 found, 48 missing, 0 empty, 0 corrupted: 100%|██████████| 5000/5000 [00:01<00:00, 2979.40it/s] -Define validation function -~~~~~~~~~~~~~~~~~~~~~~~~~~ +Define validation function `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + We will reuse validation metrics provided in the YOLOv7 repo with a modification for this case (removing extra steps). The original model @@ -691,8 +762,9 @@ Validation function reports following list of accuracy metrics: all 5000 36335 0.651 0.506 0.544 0.359 -Optimize model using NNCF Post-training Quantization API --------------------------------------------------------- +Optimize model using NNCF Post-training Quantization API `⇑ <#top>`__ +############################################################################################################################### + `NNCF `__ provides a suite of advanced algorithms for Neural Networks inference optimization in @@ -760,16 +832,30 @@ asymmetric quantization of activations. .. parsed-literal:: - Statistics collection: 100%|██████████| 300/300 [00:38<00:00, 7.87it/s] - Biases correction: 100%|██████████| 58/58 [00:04<00:00, 13.47it/s] + Statistics collection: 100%|██████████| 300/300 [00:38<00:00, 7.80it/s] + Biases correction: 100%|██████████| 58/58 [00:04<00:00, 14.15it/s] -Validate Quantized model inference ----------------------------------- +Validate Quantized model inference `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 - int8_compiled_model = core.compile_model(quantized_model, "CPU") + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + int8_compiled_model = core.compile_model(quantized_model, device.value) boxes, image, input_shape = detect(int8_compiled_model, 'inference/images/horses.jpg') image_with_boxes = draw_boxes(boxes[0], input_shape, image, NAMES, COLORS) Image.fromarray(image_with_boxes) @@ -777,12 +863,13 @@ Validate Quantized model inference -.. image:: 226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_38_0.png +.. image:: 226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_43_0.png -Validate quantized model accuracy ---------------------------------- +Validate quantized model accuracy `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -815,8 +902,8 @@ As we can see, model accuracy slightly changed after quantization. However, if we look at the output image, these changes are not significant. -Compare Performance of the Original and Quantized Models --------------------------------------------------------- +Compare Performance of the Original and Quantized Models `⇑ <#top>`__ +############################################################################################################################### Finally, use the OpenVINO `Benchmark Tool `__ @@ -830,10 +917,23 @@ models. benchmark on GPU. Run ``benchmark_app --help`` to see an overview of all command-line options. +.. code:: ipython3 + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + .. code:: ipython3 # Inference FP32 model (OpenVINO IR) - !benchmark_app -m model/yolov7-tiny.xml -d CPU -api async + !benchmark_app -m model/yolov7-tiny.xml -d $device.value -api async .. parsed-literal:: @@ -841,19 +941,20 @@ models. [Step 1/11] Parsing and validating input arguments [ INFO ] Parsing input parameters [Step 2/11] Loading OpenVINO Runtime + [ WARNING ] Default duration 120 seconds is used for unknown device AUTO [ INFO ] OpenVINO: - [ INFO ] Build ................................. 2023.0.0-10926-b4452d56304-releases/2023/0 + [ INFO ] Build ................................. 2023.0.1-11005-fa1c41994f3-releases/2023/0 [ INFO ] [ INFO ] Device info: - [ INFO ] CPU - [ INFO ] Build ................................. 2023.0.0-10926-b4452d56304-releases/2023/0 + [ INFO ] AUTO + [ INFO ] Build ................................. 2023.0.1-11005-fa1c41994f3-releases/2023/0 [ INFO ] [ INFO ] [Step 3/11] Setting device configuration - [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.THROUGHPUT. + [ WARNING ] Performance hint was not explicitly specified in command line. Device(AUTO) performance hint will be set to PerformanceMode.THROUGHPUT. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 26.87 ms + [ INFO ] Read model took 11.04 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] images (node: images) : f32 / [...] / [1,3,640,640] @@ -867,45 +968,52 @@ models. [ INFO ] Model outputs: [ INFO ] output (node: output) : f32 / [...] / [1,25200,85] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 202.20 ms + [ INFO ] Compile model took 256.97 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: + [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT [ INFO ] NETWORK_NAME: torch_jit [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 6 - [ INFO ] NUM_STREAMS: 6 - [ INFO ] AFFINITY: Affinity.CORE - [ INFO ] INFERENCE_NUM_THREADS: 24 - [ INFO ] PERF_COUNT: False - [ INFO ] INFERENCE_PRECISION_HINT: - [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT - [ INFO ] EXECUTION_MODE_HINT: ExecutionMode.PERFORMANCE - [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 - [ INFO ] ENABLE_CPU_PINNING: True - [ INFO ] SCHEDULING_CORE_TYPE: SchedulingCoreType.ANY_CORE - [ INFO ] ENABLE_HYPER_THREADING: True + [ INFO ] MODEL_PRIORITY: Priority.MEDIUM + [ INFO ] MULTI_DEVICE_PRIORITIES: CPU + [ INFO ] CPU: + [ INFO ] CPU_BIND_THREAD: YES + [ INFO ] CPU_THREADS_NUM: 0 + [ INFO ] CPU_THROUGHPUT_STREAMS: 6 + [ INFO ] DEVICE_ID: + [ INFO ] DUMP_EXEC_GRAPH_AS_DOT: + [ INFO ] DYN_BATCH_ENABLED: NO + [ INFO ] DYN_BATCH_LIMIT: 0 + [ INFO ] ENFORCE_BF16: NO + [ INFO ] EXCLUSIVE_ASYNC_REQUESTS: NO + [ INFO ] NETWORK_NAME: torch_jit + [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 6 + [ INFO ] PERFORMANCE_HINT: THROUGHPUT + [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 + [ INFO ] PERF_COUNT: NO [ INFO ] EXECUTION_DEVICES: ['CPU'] [Step 9/11] Creating infer requests and preparing input tensors [ WARNING ] No input files were given for input 'images'!. This input will be filled with random values! [ INFO ] Fill input 'images' with random values - [Step 10/11] Measuring performance (Start inference asynchronously, 6 inference requests, limits: 60000 ms duration) + [Step 10/11] Measuring performance (Start inference asynchronously, 6 inference requests, limits: 120000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 46.24 ms + [ INFO ] First inference took 43.97 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 5772 iterations - [ INFO ] Duration: 60086.63 ms + [ INFO ] Count: 11400 iterations + [ INFO ] Duration: 120097.35 ms [ INFO ] Latency: - [ INFO ] Median: 62.23 ms - [ INFO ] Average: 62.30 ms - [ INFO ] Min: 30.97 ms - [ INFO ] Max: 86.17 ms - [ INFO ] Throughput: 96.06 FPS + [ INFO ] Median: 62.78 ms + [ INFO ] Average: 63.06 ms + [ INFO ] Min: 35.00 ms + [ INFO ] Max: 133.31 ms + [ INFO ] Throughput: 94.92 FPS .. code:: ipython3 # Inference INT8 model (OpenVINO IR) - !benchmark_app -m model/yolov7-tiny_int8.xml -d CPU -api async + !benchmark_app -m model/yolov7-tiny_int8.xml -d $device.value -api async .. parsed-literal:: @@ -913,19 +1021,20 @@ models. [Step 1/11] Parsing and validating input arguments [ INFO ] Parsing input parameters [Step 2/11] Loading OpenVINO Runtime + [ WARNING ] Default duration 120 seconds is used for unknown device AUTO [ INFO ] OpenVINO: - [ INFO ] Build ................................. 2023.0.0-10926-b4452d56304-releases/2023/0 + [ INFO ] Build ................................. 2023.0.1-11005-fa1c41994f3-releases/2023/0 [ INFO ] [ INFO ] Device info: - [ INFO ] CPU - [ INFO ] Build ................................. 2023.0.0-10926-b4452d56304-releases/2023/0 + [ INFO ] AUTO + [ INFO ] Build ................................. 2023.0.1-11005-fa1c41994f3-releases/2023/0 [ INFO ] [ INFO ] [Step 3/11] Setting device configuration - [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.THROUGHPUT. + [ WARNING ] Performance hint was not explicitly specified in command line. Device(AUTO) performance hint will be set to PerformanceMode.THROUGHPUT. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 44.99 ms + [ INFO ] Read model took 17.80 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] images (node: images) : f32 / [...] / [1,3,640,640] @@ -939,37 +1048,44 @@ models. [ INFO ] Model outputs: [ INFO ] output (node: output) : f32 / [...] / [1,25200,85] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 385.78 ms + [ INFO ] Compile model took 462.21 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: + [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT [ INFO ] NETWORK_NAME: torch_jit [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 6 - [ INFO ] NUM_STREAMS: 6 - [ INFO ] AFFINITY: Affinity.CORE - [ INFO ] INFERENCE_NUM_THREADS: 24 - [ INFO ] PERF_COUNT: False - [ INFO ] INFERENCE_PRECISION_HINT: - [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT - [ INFO ] EXECUTION_MODE_HINT: ExecutionMode.PERFORMANCE - [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 - [ INFO ] ENABLE_CPU_PINNING: True - [ INFO ] SCHEDULING_CORE_TYPE: SchedulingCoreType.ANY_CORE - [ INFO ] ENABLE_HYPER_THREADING: True + [ INFO ] MODEL_PRIORITY: Priority.MEDIUM + [ INFO ] MULTI_DEVICE_PRIORITIES: CPU + [ INFO ] CPU: + [ INFO ] CPU_BIND_THREAD: YES + [ INFO ] CPU_THREADS_NUM: 0 + [ INFO ] CPU_THROUGHPUT_STREAMS: 6 + [ INFO ] DEVICE_ID: + [ INFO ] DUMP_EXEC_GRAPH_AS_DOT: + [ INFO ] DYN_BATCH_ENABLED: NO + [ INFO ] DYN_BATCH_LIMIT: 0 + [ INFO ] ENFORCE_BF16: NO + [ INFO ] EXCLUSIVE_ASYNC_REQUESTS: NO + [ INFO ] NETWORK_NAME: torch_jit + [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 6 + [ INFO ] PERFORMANCE_HINT: THROUGHPUT + [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 + [ INFO ] PERF_COUNT: NO [ INFO ] EXECUTION_DEVICES: ['CPU'] [Step 9/11] Creating infer requests and preparing input tensors [ WARNING ] No input files were given for input 'images'!. This input will be filled with random values! [ INFO ] Fill input 'images' with random values - [Step 10/11] Measuring performance (Start inference asynchronously, 6 inference requests, limits: 60000 ms duration) + [Step 10/11] Measuring performance (Start inference asynchronously, 6 inference requests, limits: 120000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 25.50 ms + [ INFO ] First inference took 26.85 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 15864 iterations - [ INFO ] Duration: 60030.26 ms + [ INFO ] Count: 31326 iterations + [ INFO ] Duration: 120015.35 ms [ INFO ] Latency: - [ INFO ] Median: 22.53 ms - [ INFO ] Average: 22.58 ms - [ INFO ] Min: 12.63 ms - [ INFO ] Max: 42.07 ms - [ INFO ] Throughput: 264.27 FPS + [ INFO ] Median: 22.78 ms + [ INFO ] Average: 22.86 ms + [ INFO ] Min: 14.12 ms + [ INFO ] Max: 41.51 ms + [ INFO ] Throughput: 261.02 FPS diff --git a/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_22_0.jpg b/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_22_0.jpg deleted file mode 100644 index df63e91b95c..00000000000 --- a/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_22_0.jpg +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:c88a66bcb937be53feb08be775b268db52543cafc343ed6ff0a3fca0ed6f97a5 -size 64211 diff --git a/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_22_0.png b/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_22_0.png deleted file mode 100644 index f0aa34d79cd..00000000000 --- a/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_22_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:e89095ad37719982e3ba5e3e106d7b5499d9412a78e757fe21e8558661ec8b87 -size 574062 diff --git a/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_26_0.jpg b/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_26_0.jpg new file mode 100644 index 00000000000..233197348b4 --- /dev/null +++ b/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_26_0.jpg @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:2975c3827089d20aad73d540189b9fd908f30dfeff795ff14e68e95327d24d35 +size 63563 diff --git a/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_26_0.png b/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_26_0.png new file mode 100644 index 00000000000..b878c1c054b --- /dev/null +++ b/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_26_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:f397bd82701c329f4e73ff181b5d1c4ac907f4b61910ca09681f63ff83659875 +size 573281 diff --git a/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_38_0.jpg b/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_38_0.jpg deleted file mode 100644 index cf23cd88049..00000000000 --- a/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_38_0.jpg +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:0fc5748c620e7cdcd55692d46d6a4e978482d44867355a51903dd8ea59a1e250 -size 64202 diff --git a/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_38_0.png b/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_38_0.png deleted file mode 100644 index 40e605f0599..00000000000 --- a/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_38_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:9cb6d21efa0da3cb90834e8dc748ff9f2b4944c722290bb667c872f3e295bebf -size 573420 diff --git a/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_43_0.jpg b/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_43_0.jpg new file mode 100644 index 00000000000..9f6428a8dbc --- /dev/null +++ b/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_43_0.jpg @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:1513cf0169b7ae87be5234830d8fbaaf3a766a08d9e597bd75ff55ea6d0e2892 +size 63483 diff --git a/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_43_0.png b/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_43_0.png new file mode 100644 index 00000000000..a807e7cfdd1 --- /dev/null +++ b/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_43_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:cd195b83150860a64d32e8140921e2b1494ed619b8c65789e46fe5ba6d9e4802 +size 572961 diff --git a/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_8_0.jpg b/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_8_0.jpg deleted file mode 100644 index 32a86187864..00000000000 --- a/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_8_0.jpg +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:37b0d33ceb3ab743bc4739296fe5973b6c5a9231c9547fc68b26fe4ec24c3f5e -size 63799 diff --git a/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_8_0.png b/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_8_0.png deleted file mode 100644 index 5b66b727c19..00000000000 --- a/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_8_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:560ecfb28dec4a920d8843cbe843d36e0c2899ed2bf039343e12b61b806a5172 -size 566136 diff --git a/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_9_0.jpg b/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_9_0.jpg new file mode 100644 index 00000000000..62f74cb495e --- /dev/null +++ b/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_9_0.jpg @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:a8f7a6a17776622d2117a24a02238d0062fc727c9b093a25b6f5c77584416cf0 +size 64007 diff --git a/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_9_0.png b/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_9_0.png new file mode 100644 index 00000000000..b0d6bbcdae9 --- /dev/null +++ b/docs/notebooks/226-yolov7-optimization-with-output_files/226-yolov7-optimization-with-output_9_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:77cb641bd57ccfe7c494b2e7b556db50ff5fb0f4c3d2d999ec94fe7679fc3919 +size 568350 diff --git a/docs/notebooks/226-yolov7-optimization-with-output_files/index.html b/docs/notebooks/226-yolov7-optimization-with-output_files/index.html index 1a6335a6186..a0f9f8664b5 100644 --- a/docs/notebooks/226-yolov7-optimization-with-output_files/index.html +++ b/docs/notebooks/226-yolov7-optimization-with-output_files/index.html @@ -1,12 +1,12 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/226-yolov7-optimization-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/226-yolov7-optimization-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/226-yolov7-optimization-with-output_files/


../
-226-yolov7-optimization-with-output_22_0.jpg       12-Jul-2023 00:11               64211
-226-yolov7-optimization-with-output_22_0.png       12-Jul-2023 00:11              574062
-226-yolov7-optimization-with-output_38_0.jpg       12-Jul-2023 00:11               64202
-226-yolov7-optimization-with-output_38_0.png       12-Jul-2023 00:11              573420
-226-yolov7-optimization-with-output_8_0.jpg        12-Jul-2023 00:11               63799
-226-yolov7-optimization-with-output_8_0.png        12-Jul-2023 00:11              566136
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/226-yolov7-optimization-with-output_files/


../
+226-yolov7-optimization-with-output_26_0.jpg       16-Aug-2023 01:31               63563
+226-yolov7-optimization-with-output_26_0.png       16-Aug-2023 01:31              573281
+226-yolov7-optimization-with-output_43_0.jpg       16-Aug-2023 01:31               63483
+226-yolov7-optimization-with-output_43_0.png       16-Aug-2023 01:31              572961
+226-yolov7-optimization-with-output_9_0.jpg        16-Aug-2023 01:31               64007
+226-yolov7-optimization-with-output_9_0.png        16-Aug-2023 01:31              568350
 

diff --git a/docs/notebooks/227-whisper-subtitles-generation-with-output.rst b/docs/notebooks/227-whisper-subtitles-generation-with-output.rst index 74bcae9556c..05b04c2fec8 100644 --- a/docs/notebooks/227-whisper-subtitles-generation-with-output.rst +++ b/docs/notebooks/227-whisper-subtitles-generation-with-output.rst @@ -1,6 +1,8 @@ Video Subtitle Generation using Whisper and OpenVINO™ ===================================================== +.. _top: + `Whisper `__ is an automatic speech recognition (ASR) system trained on 680,000 hours of multilingual and multitask supervised data collected from the web. It is a multi-task @@ -21,26 +23,47 @@ GitHub `repository `__. In this notebook, we will use Whisper with OpenVINO to generate subtitles in a sample video. Notebook contains the following steps: 1. Download the model. 2. Instantiate the PyTorch model pipeline. 3. Export -the ONNX model and convert it to OpenVINO IR, using the Model Optimizer -tool. 4. Run the Whisper pipeline with OpenVINO models. +the ONNX model and convert it to OpenVINO IR, using model conversion +API. 4. Run the Whisper pipeline with OpenVINO models. + +**Table of contents**: + +- `Prerequisites <#prerequisites>`__ +- `Instantiate model <#instantiate-model>`__ + + - `Convert model to OpenVINO Intermediate Representation (IR) format. <#convert-model-to-openvino-intermediate-representation-ir-format>`__ + - `Convert Whisper Encoder to OpenVINO IR <#convert-whisper-encoder-to-openvino-ir>`__ + - `Convert Whisper decoder to OpenVINO IR <#5convert-whisper-decoder-to-openvino-ir>`__ + +- `Prepare inference pipeline <#prepare-inference-pipeline>`__ + + - `Select inference device <#select-inference-device>`__ + + - `Define audio preprocessing <#define-audio-preprocessing>`__ + +- `Run video transcription pipeline <#run-video-transcription-pipeline>`__ + +Prerequisites `⇑ <#top>`__ +############################################################################################################################### -Prerequisites -------------- Clone and install the model repository. .. code:: ipython3 - !pip install -q 'openvino-dev>=2023.0.0' + !pip install -q "openvino-dev>=2023.0.0" !pip install -q "python-ffmpeg<=1.0.16" moviepy transformers onnx !pip install -q -I "git+https://github.com/garywu007/pytube.git" .. parsed-literal:: + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 ERROR: pip's dependency resolver does not currently take into account all the packages that are installed. This behaviour is the source of the following dependency conflicts. ppgan 2.1.0 requires librosa==0.8.1, but you have librosa 0.9.2 which is incompatible. - ppgan 2.1.0 requires opencv-python<=4.6.0.66, but you have opencv-python 4.8.0.74 which is incompatible. + ppgan 2.1.0 requires opencv-python<=4.6.0.66, but you have opencv-python 4.8.0.76 which is incompatible. + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 .. code:: ipython3 @@ -56,12 +79,12 @@ Clone and install the model repository. .. parsed-literal:: Cloning into 'whisper'... - remote: Enumerating objects: 585, done. - remote: Counting objects: 100% (304/304), done. - remote: Compressing objects: 100% (53/53), done. - remote: Total 585 (delta 275), reused 253 (delta 251), pack-reused 281 - Receiving objects: 100% (585/585), 8.14 MiB | 3.85 MiB/s, done. - Resolving deltas: 100% (352/352), done. + remote: Enumerating objects: 589, done. + remote: Counting objects: 100% (367/367), done. + remote: Compressing objects: 100% (82/82), done. + remote: Total 589 (delta 320), reused 288 (delta 285), pack-reused 222 + Receiving objects: 100% (589/589), 8.14 MiB | 4.18 MiB/s, done. + Resolving deltas: 100% (357/357), done. Note: switching to '55f690af7914c672c69733b7e04ef5a41b2b2774'. You are in 'detached HEAD' state. You can look around, make experimental @@ -79,51 +102,53 @@ Clone and install the model repository. Turn off this advice by setting config variable advice.detachedHead to false - Processing /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/227-whisper-subtitles-generation/whisper + Processing /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/227-whisper-subtitles-generation/whisper Preparing metadata (setup.py) ... - done - Requirement already satisfied: numpy in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from openai-whisper==20230124) (1.23.5) - Requirement already satisfied: torch in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from openai-whisper==20230124) (1.13.1+cpu) - Requirement already satisfied: tqdm in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from openai-whisper==20230124) (4.65.0) + Requirement already satisfied: numpy in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from openai-whisper==20230124) (1.23.5) + Requirement already satisfied: torch in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from openai-whisper==20230124) (1.13.1+cpu) + Requirement already satisfied: tqdm in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from openai-whisper==20230124) (4.66.1) Collecting more-itertools (from openai-whisper==20230124) - Using cached more_itertools-9.1.0-py3-none-any.whl (54 kB) - Requirement already satisfied: transformers>=4.19.0 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from openai-whisper==20230124) (4.30.2) + Obtaining dependency information for more-itertools from https://files.pythonhosted.org/packages/5a/cb/6dce742ea14e47d6f565589e859ad225f2a5de576d7696e0623b784e226b/more_itertools-10.1.0-py3-none-any.whl.metadata + Using cached more_itertools-10.1.0-py3-none-any.whl.metadata (33 kB) + Requirement already satisfied: transformers>=4.19.0 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from openai-whisper==20230124) (4.31.0) Collecting ffmpeg-python==0.2.0 (from openai-whisper==20230124) Using cached ffmpeg_python-0.2.0-py3-none-any.whl (25 kB) - Requirement already satisfied: future in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ffmpeg-python==0.2.0->openai-whisper==20230124) (0.18.3) - Requirement already satisfied: filelock in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.19.0->openai-whisper==20230124) (3.12.2) - Requirement already satisfied: huggingface-hub<1.0,>=0.14.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.19.0->openai-whisper==20230124) (0.16.4) - Requirement already satisfied: packaging>=20.0 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.19.0->openai-whisper==20230124) (23.1) - Requirement already satisfied: pyyaml>=5.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.19.0->openai-whisper==20230124) (6.0) - Requirement already satisfied: regex!=2019.12.17 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.19.0->openai-whisper==20230124) (2023.6.3) - Requirement already satisfied: requests in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.19.0->openai-whisper==20230124) (2.31.0) - Requirement already satisfied: tokenizers!=0.11.3,<0.14,>=0.11.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.19.0->openai-whisper==20230124) (0.13.3) - Requirement already satisfied: safetensors>=0.3.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.19.0->openai-whisper==20230124) (0.3.1) - Requirement already satisfied: typing-extensions in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from torch->openai-whisper==20230124) (4.7.1) - Requirement already satisfied: fsspec in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from huggingface-hub<1.0,>=0.14.1->transformers>=4.19.0->openai-whisper==20230124) (2023.6.0) - Requirement already satisfied: charset-normalizer<4,>=2 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests->transformers>=4.19.0->openai-whisper==20230124) (3.2.0) - Requirement already satisfied: idna<4,>=2.5 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests->transformers>=4.19.0->openai-whisper==20230124) (3.4) - Requirement already satisfied: urllib3<3,>=1.21.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests->transformers>=4.19.0->openai-whisper==20230124) (1.26.16) - Requirement already satisfied: certifi>=2017.4.17 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests->transformers>=4.19.0->openai-whisper==20230124) (2023.5.7) + Requirement already satisfied: future in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ffmpeg-python==0.2.0->openai-whisper==20230124) (0.18.3) + Requirement already satisfied: filelock in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.19.0->openai-whisper==20230124) (3.12.2) + Requirement already satisfied: huggingface-hub<1.0,>=0.14.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.19.0->openai-whisper==20230124) (0.16.4) + Requirement already satisfied: packaging>=20.0 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.19.0->openai-whisper==20230124) (23.1) + Requirement already satisfied: pyyaml>=5.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.19.0->openai-whisper==20230124) (6.0.1) + Requirement already satisfied: regex!=2019.12.17 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.19.0->openai-whisper==20230124) (2023.8.8) + Requirement already satisfied: requests in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.19.0->openai-whisper==20230124) (2.31.0) + Requirement already satisfied: tokenizers!=0.11.3,<0.14,>=0.11.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.19.0->openai-whisper==20230124) (0.13.3) + Requirement already satisfied: safetensors>=0.3.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.19.0->openai-whisper==20230124) (0.3.2) + Requirement already satisfied: typing-extensions in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from torch->openai-whisper==20230124) (4.7.1) + Requirement already satisfied: fsspec in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from huggingface-hub<1.0,>=0.14.1->transformers>=4.19.0->openai-whisper==20230124) (2023.6.0) + Requirement already satisfied: charset-normalizer<4,>=2 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests->transformers>=4.19.0->openai-whisper==20230124) (3.2.0) + Requirement already satisfied: idna<4,>=2.5 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests->transformers>=4.19.0->openai-whisper==20230124) (3.4) + Requirement already satisfied: urllib3<3,>=1.21.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests->transformers>=4.19.0->openai-whisper==20230124) (1.26.16) + Requirement already satisfied: certifi>=2017.4.17 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests->transformers>=4.19.0->openai-whisper==20230124) (2023.7.22) + Using cached more_itertools-10.1.0-py3-none-any.whl (55 kB) Building wheels for collected packages: openai-whisper Building wheel for openai-whisper (setup.py) ... - \ | done - Created wheel for openai-whisper: filename=openai_whisper-20230124-py3-none-any.whl size=1179311 sha256=a75d8198e0a3343e35889bda79a05e0b83878dae0f191cba9fd4ba44f5ca6ae6 - Stored in directory: /tmp/pip-ephem-wheel-cache-4_3s8aog/wheels/07/04/88/d2c2d6f2253db7e60a09770f6870703d1fd581296886334a97 + Created wheel for openai-whisper: filename=openai_whisper-20230124-py3-none-any.whl size=1179305 sha256=4fcfbe9ab46c8d5e7a7fa0c52e896e59bdbc043a743c686acc001c6ed8dc5e65 + Stored in directory: /tmp/pip-ephem-wheel-cache-5a4nqoja/wheels/0c/9d/b6/d90fb003a36a5e4026f7e998e937791cc6a6c6e9abea61d48d Successfully built openai-whisper + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 Installing collected packages: more-itertools, ffmpeg-python, openai-whisper - Successfully installed ffmpeg-python-0.2.0 more-itertools-9.1.0 openai-whisper-20230124 + Successfully installed ffmpeg-python-0.2.0 more-itertools-10.1.0 openai-whisper-20230124 -Instantiate model ------------------ +Instantiate model `⇑ <#top>`__ +############################################################################################################################### -Whisper is a Transformer based encoder-decoder model, also referred to -as a sequence-to-sequence model. It maps a sequence of audio spectrogram -features to a sequence of text tokens. First, the raw audio inputs are -converted to a log-Mel spectrogram by action of the feature extractor. -Then, the Transformer encoder encodes the spectrogram to form a sequence -of encoder hidden states. Finally, the decoder autoregressively predicts -text tokens, conditional on both the previous tokens and the encoder -hidden states. +Whisper is a Transformer based encoder-decoder model, also referred to as a sequence-to-sequence model. +It maps a sequence of audio spectrogram features to a sequence of text +tokens. First, the raw audio inputs are converted to a log-Mel +spectrogram by action of the feature extractor. Then, the Transformer +encoder encodes the spectrogram to form a sequence of encoder hidden +states. Finally, the decoder autoregressively predicts text tokens, +conditional on both the previous tokens and the encoder hidden states. You can see the model architecture in the diagram below: @@ -146,8 +171,8 @@ Whisper family. model.eval() pass -Convert model to OpenVINO Intermediate Representation (IR) format. -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Convert model to OpenVINO Intermediate Representation (IR) format. `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ For best results with OpenVINO, it is recommended to convert the model to OpenVINO IR format. OpenVINO supports PyTorch via ONNX conversion. We @@ -159,8 +184,9 @@ Python function returns an OpenVINO model ready to load on device and start making predictions. We can save it on disk for next usage with ``openvino.runtime.serialize``. -Convert Whisper Encoder to OpenVINO IR -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Convert Whisper Encoder to OpenVINO IR `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -183,12 +209,13 @@ Convert Whisper Encoder to OpenVINO IR .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/whisper/model.py:153: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/whisper/model.py:153: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! assert x.shape[1:] == self.positional_embedding.shape, "incorrect audio shape" -Convert Whisper decoder to OpenVINO IR -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Convert Whisper decoder to OpenVINO IR `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + To reduce computational complexity, the decoder uses cached key/value projections in attention modules from the previous steps. We need to @@ -362,7 +389,7 @@ modify this process for correct tracing to ONNX. .. parsed-literal:: - /tmp/ipykernel_3462814/1737529362.py:18: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /tmp/ipykernel_2070841/1737529362.py:18: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if module not in cache or output.shape[1] > positional_embeddings_size: @@ -375,18 +402,19 @@ input shapes. .. code:: ipython3 - input_shapes = "tokens[1..5 1..224],audio_features[1..5 1500 512]" + input_shapes = "tokens[1..5 -1],audio_features[1..5 1500 512]" for k, v in kv_cache.items(): if k.endswith('a'): - input_shapes += f",in_{k}[1..5 0..224 512]" + input_shapes += f",in_{k}[1..5 -1 512]" decoder_model = mo.convert_model( input_model="whisper_decoder.onnx", compress_to_fp16=True, input=input_shapes) serialize(decoder_model, "whisper_decoder.xml") -Prepare inference pipeline --------------------------- +Prepare inference pipeline `⇑ <#top>`__ +############################################################################################################################### + The image below illustrates the pipeline of video transcribing using the Whisper model. @@ -617,16 +645,46 @@ original models with OpenVINO IR versions. del model.decoder del model.encoder +.. code:: ipython3 + + core = Core() + +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + .. code:: ipython3 from collections import namedtuple Parameter = namedtuple('Parameter', ['device']) - core = Core() - - model.encoder = OpenVINOAudioEncoder(core, 'whisper_encoder.xml') - model.decoder = OpenVINOTextDecoder(core, 'whisper_decoder.xml') + model.encoder = OpenVINOAudioEncoder(core, 'whisper_encoder.xml', device=device.value) + model.decoder = OpenVINOTextDecoder(core, 'whisper_decoder.xml', device=device.value) model.decode = partial(decode, model) @@ -651,8 +709,9 @@ original models with OpenVINO IR versions. model.logits = partial(logits, model) -Define audio preprocessing -^^^^^^^^^^^^^^^^^^^^^^^^^^ +Define audio preprocessing `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + The model expects mono-channel audio with a 16000 Hz sample rate, represented in floating point range. When the audio from the input video @@ -718,8 +777,9 @@ does not meet these requirements, we will need to apply preprocessing. resampled_audio = resample(audio, sample_rate, 16000) return resampled_audio -Run video transcription pipeline --------------------------------- +Run video transcription pipeline `⇑ <#top>`__ +############################################################################################################################### + Now, we are ready to start transcription. We select a video from YouTube that we want to transcribe. Be patient, as downloading the video may @@ -769,8 +829,10 @@ take some time. Select the task for the model: -* **transcribe** - generate audio transcription in the source language (automatically detected). -* **translate** - generate audio transcription with translation to English language. +- **transcribe** - generate audio transcription in the source language + (automatically detected). +- **translate** - generate audio transcription with translation to + English language. .. code:: ipython3 diff --git a/docs/notebooks/228-clip-zero-shot-image-classification-with-output.rst b/docs/notebooks/228-clip-zero-shot-convert-with-output.rst similarity index 62% rename from docs/notebooks/228-clip-zero-shot-image-classification-with-output.rst rename to docs/notebooks/228-clip-zero-shot-convert-with-output.rst index d4e0ae21c41..913817a8a4e 100644 --- a/docs/notebooks/228-clip-zero-shot-image-classification-with-output.rst +++ b/docs/notebooks/228-clip-zero-shot-convert-with-output.rst @@ -1,6 +1,8 @@ Zero-shot Image Classification with OpenAI CLIP and OpenVINO™ ============================================================= +.. _top: + Zero-shot image classification is a computer vision task to classify images into one of several classes without any prior training or knowledge of the classes. @@ -20,18 +22,30 @@ associate unseen categories to images with zero-shot learning by exploiting attributes to model’s relationship between visual features and labels. In this tutorial, we will use the `OpenAI CLIP `__ model to perform zero-shot -image classification. +image classification. The notebook contains the following steps: -The notebook contains the following steps: - -1. Download the model. -2. Instantiate the PyTorch model. -3. Export the ONNX -model and convert it to OpenVINO IR, using the Model Optimizer tool. +1. Download the model. +2. Instantiate the PyTorch model. +3. Export the ONNX model and convert it to OpenVINO IR, using model + conversion API. 4. Run CLIP with OpenVINO. -Instantiate model ------------------ +**Table of contents**: + +- `Instantiate model <#instantiate-model>`__ +- `Run PyTorch model inference <#run-pytorch-model-inference>`__ + + - `Convert model to OpenVINO Intermediate Representation (IR) format. <#convert-model-to-openvino-intermediate-representation-ir-format>`__ + +- `Run OpenVINO model <#run-openvino-model>`__ + + - `Select inference device <#select-inference-device>`__ + +- `Next Steps <#next-steps>`__ + +Instantiate model `⇑ <#top>`__ +############################################################################################################################### + CLIP (Contrastive Language-Image Pre-Training) is a neural network trained on various (image, text) pairs. It can be instructed in natural @@ -78,51 +92,57 @@ tokenizer and preparing the images. processor = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch16") + .. parsed-literal:: - 2023-07-11 23:27:00.851579: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. - 2023-07-11 23:27:00.884422: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. - To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. - 2023-07-11 23:27:01.353735: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + Downloading (…)lve/main/config.json: 0%| | 0.00/4.10k [00:00`__ +############################################################################################################################### -Run PyTorch model inference ---------------------------- To perform classification, define labels and load an image in RGB format. To give the model wider text context and improve guidance, we @@ -136,6 +156,9 @@ similarity score for the final result. .. code:: ipython3 + from PIL import Image + from visualize import visualize_result + image = Image.open('../data/image/coco.jpg') input_labels = ['cat', 'dog', 'wolf', 'tiger', 'man', 'horse', 'frog', 'tree', 'house', 'computer'] text_descriptions = [f"This is a photo of a {label}" for label in input_labels] @@ -149,11 +172,16 @@ similarity score for the final result. -.. image:: 228-clip-zero-shot-image-classification-with-output_files/228-clip-zero-shot-image-classification-with-output_5_0.png +.. image:: 228-clip-zero-shot-convert-with-output_files/228-clip-zero-shot-convert-with-output_4_0.png -Convert model to OpenVINO Intermediate Representation (IR) format. -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Convert model to OpenVINO Intermediate Representation (IR) format. `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + +.. figure:: https://user-images.githubusercontent.com/29454499/208048580-8264e54c-151c-43ef-9e25-1302cd0dd7a2.png + :alt: conversion_path + + conversion_path For best results with OpenVINO, it is recommended to convert the model to OpenVINO IR format. OpenVINO supports PyTorch via ONNX conversion. @@ -200,17 +228,17 @@ for the next usage with ``openvino.runtime.serialize``. .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/clip/modeling_clip.py:284: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /home/adrian/repos/openvino_notebooks/recipes/intelligent_queue_management/venv/lib/python3.10/site-packages/transformers/models/clip/modeling_clip.py:284: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if attn_weights.size() != (bsz * self.num_heads, tgt_len, src_len): - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/clip/modeling_clip.py:324: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /home/adrian/repos/openvino_notebooks/recipes/intelligent_queue_management/venv/lib/python3.10/site-packages/transformers/models/clip/modeling_clip.py:324: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if attn_output.size() != (bsz * self.num_heads, tgt_len, self.head_dim): - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/clip/modeling_clip.py:684: TracerWarning: torch.tensor results are registered as constants in the trace. You can safely ignore this warning if you use this function to create tensors out of constant variables that would be the same every time you call this function. In any other case, this might cause the trace to be incorrect. + /home/adrian/repos/openvino_notebooks/recipes/intelligent_queue_management/venv/lib/python3.10/site-packages/transformers/models/clip/modeling_clip.py:684: TracerWarning: torch.tensor results are registered as constants in the trace. You can safely ignore this warning if you use this function to create tensors out of constant variables that would be the same every time you call this function. In any other case, this might cause the trace to be incorrect. mask = torch.full((tgt_len, tgt_len), torch.tensor(torch.finfo(dtype).min, device=device), device=device) - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/clip/modeling_clip.py:292: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /home/adrian/repos/openvino_notebooks/recipes/intelligent_queue_management/venv/lib/python3.10/site-packages/transformers/models/clip/modeling_clip.py:292: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if causal_attention_mask.size() != (bsz, 1, tgt_len, src_len): - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/clip/modeling_clip.py:301: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /home/adrian/repos/openvino_notebooks/recipes/intelligent_queue_management/venv/lib/python3.10/site-packages/transformers/models/clip/modeling_clip.py:301: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if attention_mask.size() != (bsz, 1, tgt_len, src_len): - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/symbolic_opset9.py:5408: UserWarning: Exporting aten::index operator of advanced indexing in opset 14 is achieved by combination of multiple ONNX operators, including Reshape, Transpose, Concat, and Gather. If indices include negative values, the exported graph will produce incorrect results. + /home/adrian/repos/openvino_notebooks/recipes/intelligent_queue_management/venv/lib/python3.10/site-packages/torch/onnx/symbolic_opset9.py:5408: UserWarning: Exporting aten::index operator of advanced indexing in opset 14 is achieved by combination of multiple ONNX operators, including Reshape, Transpose, Concat, and Gather. If indices include negative values, the exported graph will produce incorrect results. warnings.warn( @@ -222,17 +250,9 @@ for the next usage with ``openvino.runtime.serialize``. ov_model = mo.convert_model('clip-vit-base-patch16.onnx', compress_to_fp16=True) serialize(ov_model, 'clip-vit-base-patch16.xml') +Run OpenVINO model `⇑ <#top>`__ +############################################################################################################################### -.. parsed-literal:: - - huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks... - To disable this warning, you can either: - - Avoid using `tokenizers` before the fork if possible - - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false) - - -Run OpenVINO model ------------------- The steps for making predictions with the OpenVINO CLIP model are similar to the PyTorch model. Let us check the model result using the @@ -240,14 +260,44 @@ same input data from the example above with PyTorch. .. code:: ipython3 - import numpy as np from scipy.special import softmax from openvino.runtime import Core # create OpenVINO core object instance core = Core() + +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=3, options=('CPU', 'GPU.0', 'GPU.1', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + # compile model for loading on device - compiled_model = core.compile_model(ov_model) + compiled_model = core.compile_model(ov_model, device.value) # obtain output tensor for getting predictions logits_per_image_out = compiled_model.output(0) # run inference on preprocessed data and get image-text similarity score @@ -259,7 +309,7 @@ same input data from the example above with PyTorch. -.. image:: 228-clip-zero-shot-image-classification-with-output_files/228-clip-zero-shot-image-classification-with-output_10_0.png +.. image:: 228-clip-zero-shot-convert-with-output_files/228-clip-zero-shot-convert-with-output_12_0.png Great! Looks like we got the same result. @@ -324,5 +374,14 @@ Run the next cell to get the result for your submitted data: -.. image:: 228-clip-zero-shot-image-classification-with-output_files/228-clip-zero-shot-image-classification-with-output_15_0.png +.. image:: 228-clip-zero-shot-convert-with-output_files/228-clip-zero-shot-convert-with-output_17_0.png + +Next Steps `⇑ <#top>`__ +############################################################################################################################### + + +Open the +`228-clip-zero-shot-quantize <228-clip-zero-shot-quantize.ipynb>`__ +notebook to quantize the IR model with the Post-training Quantization +API of NNCF and compare ``FP16`` and ``INT8`` models. diff --git a/docs/notebooks/228-clip-zero-shot-convert-with-output_files/228-clip-zero-shot-convert-with-output_12_0.png b/docs/notebooks/228-clip-zero-shot-convert-with-output_files/228-clip-zero-shot-convert-with-output_12_0.png new file mode 100644 index 00000000000..e8137a5ea46 --- /dev/null +++ b/docs/notebooks/228-clip-zero-shot-convert-with-output_files/228-clip-zero-shot-convert-with-output_12_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:c7b8ae280413c012c0cc0c2c4df95fe57b56a971584a536d32a46442b9d89c4b +size 464100 diff --git a/docs/notebooks/228-clip-zero-shot-convert-with-output_files/228-clip-zero-shot-convert-with-output_17_0.png b/docs/notebooks/228-clip-zero-shot-convert-with-output_files/228-clip-zero-shot-convert-with-output_17_0.png new file mode 100644 index 00000000000..2f089a8dbec --- /dev/null +++ b/docs/notebooks/228-clip-zero-shot-convert-with-output_files/228-clip-zero-shot-convert-with-output_17_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:c7a20c293a356fb88fdd3713c032b5093b6acf2d03ab4c98a70b11829d1c75ab +size 461829 diff --git a/docs/notebooks/228-clip-zero-shot-convert-with-output_files/228-clip-zero-shot-convert-with-output_4_0.png b/docs/notebooks/228-clip-zero-shot-convert-with-output_files/228-clip-zero-shot-convert-with-output_4_0.png new file mode 100644 index 00000000000..e8137a5ea46 --- /dev/null +++ b/docs/notebooks/228-clip-zero-shot-convert-with-output_files/228-clip-zero-shot-convert-with-output_4_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:c7b8ae280413c012c0cc0c2c4df95fe57b56a971584a536d32a46442b9d89c4b +size 464100 diff --git a/docs/notebooks/228-clip-zero-shot-convert-with-output_files/index.html b/docs/notebooks/228-clip-zero-shot-convert-with-output_files/index.html new file mode 100644 index 00000000000..30057e0e2eb --- /dev/null +++ b/docs/notebooks/228-clip-zero-shot-convert-with-output_files/index.html @@ -0,0 +1,9 @@ + +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/228-clip-zero-shot-convert-with-output_files/ + +

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/228-clip-zero-shot-convert-with-output_files/


../
+228-clip-zero-shot-convert-with-output_12_0.png    16-Aug-2023 01:31              464100
+228-clip-zero-shot-convert-with-output_17_0.png    16-Aug-2023 01:31              461829
+228-clip-zero-shot-convert-with-output_4_0.png     16-Aug-2023 01:31              464100
+

+ diff --git a/docs/notebooks/228-clip-zero-shot-image-classification-with-output_files/228-clip-zero-shot-image-classification-with-output_10_0.png b/docs/notebooks/228-clip-zero-shot-image-classification-with-output_files/228-clip-zero-shot-image-classification-with-output_10_0.png deleted file mode 100644 index 64d795fb305..00000000000 --- a/docs/notebooks/228-clip-zero-shot-image-classification-with-output_files/228-clip-zero-shot-image-classification-with-output_10_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:1349e0c8dcd14884b2111ff68adcabba70fb4be84d3e7bd4c854c485f765288b -size 464100 diff --git a/docs/notebooks/228-clip-zero-shot-image-classification-with-output_files/228-clip-zero-shot-image-classification-with-output_15_0.png b/docs/notebooks/228-clip-zero-shot-image-classification-with-output_files/228-clip-zero-shot-image-classification-with-output_15_0.png deleted file mode 100644 index 7b8b1c784a4..00000000000 --- a/docs/notebooks/228-clip-zero-shot-image-classification-with-output_files/228-clip-zero-shot-image-classification-with-output_15_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:8e069175c9e618fb1d796b72743a56f9ecbdaf5adfaf4e9e8c56ba7ab6428cd8 -size 461829 diff --git a/docs/notebooks/228-clip-zero-shot-image-classification-with-output_files/228-clip-zero-shot-image-classification-with-output_5_0.png b/docs/notebooks/228-clip-zero-shot-image-classification-with-output_files/228-clip-zero-shot-image-classification-with-output_5_0.png deleted file mode 100644 index 64d795fb305..00000000000 --- a/docs/notebooks/228-clip-zero-shot-image-classification-with-output_files/228-clip-zero-shot-image-classification-with-output_5_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:1349e0c8dcd14884b2111ff68adcabba70fb4be84d3e7bd4c854c485f765288b -size 464100 diff --git a/docs/notebooks/228-clip-zero-shot-image-classification-with-output_files/index.html b/docs/notebooks/228-clip-zero-shot-image-classification-with-output_files/index.html deleted file mode 100644 index 6d03f41cb71..00000000000 --- a/docs/notebooks/228-clip-zero-shot-image-classification-with-output_files/index.html +++ /dev/null @@ -1,9 +0,0 @@ - -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/228-clip-zero-shot-image-classification-with-output_files/ - -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/228-clip-zero-shot-image-classification-with-output_files/


../
-228-clip-zero-shot-image-classification-with-ou..> 12-Jul-2023 00:11              464100
-228-clip-zero-shot-image-classification-with-ou..> 12-Jul-2023 00:11              461829
-228-clip-zero-shot-image-classification-with-ou..> 12-Jul-2023 00:11              464100
-

- diff --git a/docs/notebooks/228-clip-zero-shot-quantize-with-output.rst b/docs/notebooks/228-clip-zero-shot-quantize-with-output.rst new file mode 100644 index 00000000000..9414a38bf56 --- /dev/null +++ b/docs/notebooks/228-clip-zero-shot-quantize-with-output.rst @@ -0,0 +1,376 @@ +Post-Training Quantization of OpenAI CLIP model with NNCF +========================================================= + +.. _top: + +The goal of this tutorial is to demonstrate how to speed up the model by +applying 8-bit post-training quantization from +`NNCF `__ (Neural Network +Compression Framework) and infer quantized model via OpenVINO™ Toolkit. +The optimization process contains the following steps: + +1. Quantize the converted OpenVINO model from + `notebook <228-clip-zero-shot-convert.ipynb>`__ with NNCF. +2. Check the model result using the same input data from the + `notebook <228-clip-zero-shot-convert.ipynb>`__. +3. Compare model size of converted and quantized models. +4. Compare performance of converted and quantized models. + +.. + + **NOTE**: you should run + `228-clip-zero-shot-convert <228-clip-zero-shot-convert.ipynb>`__ + notebook first to generate OpenVINO IR model that is used for + quantization. + +**Table of contents**: + +- `Prerequisites <#prerequisites>`__ +- `Create and initialize quantization <#create-and-initialize-quantization>`__ + + - `Prepare datasets <#prepare-datasets>`__ + +- `Run quantized OpenVINO model <#run-quantized-openvino-model>`__ + + - `Compare File Size <#compare-file-size>`__ + - `Compare inference time of the FP16 IR and quantized models <#compare-inference-time-of-the-fp16-ir-and-quantized-models>`__ + +Prerequisites `⇑ <#top>`__ +############################################################################################################################### + + +.. code:: ipython3 + + !pip install -q datasets + !pip install -q "git+https://github.com/openvinotoolkit/nncf.git@6c0aebadd2fcdbe1481a11b40b8cd9f66b3b6fab" + +Create and initialize quantization `⇑ <#top>`__ +############################################################################################################################### + + +`NNCF `__ enables +post-training quantization by adding the quantization layers into the +model graph and then using a subset of the training dataset to +initialize the parameters of these additional quantization layers. The +framework is designed so that modifications to your original training +code are minor. Quantization is the simplest scenario and requires a few +modifications. + +The optimization process contains the following steps: + +1. Create a Dataset for quantization. +2. Run ``nncf.quantize`` for getting a quantized model. +3. Serialize the ``INT8`` model using ``openvino.runtime.serialize`` + function. + +Prepare datasets `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +The `Conceptual +Captions `__ dataset +consisting of ~3.3M images annotated with captions is used to quantize +model. + +.. code:: ipython3 + + import os + + fp16_model_path = 'clip-vit-base-patch16.xml' + if not os.path.exists(fp16_model_path): + raise RuntimeError('This notebook should be run after 228-clip-zero-shot-convert.ipynb.') + +.. code:: ipython3 + + from transformers import CLIPProcessor, CLIPModel + + model = CLIPModel.from_pretrained("openai/clip-vit-base-patch16") + max_length = model.config.text_config.max_position_embeddings + processor = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch16") + +.. code:: ipython3 + + import requests + from io import BytesIO + from PIL import Image + from requests.packages.urllib3.exceptions import InsecureRequestWarning + requests.packages.urllib3.disable_warnings(InsecureRequestWarning) + + def check_text_data(data): + """ + Check if the given data is text-based. + """ + if isinstance(data, str): + return True + if isinstance(data, list): + return all(isinstance(x, str) for x in data) + return False + + def get_pil_from_url(url): + """ + Downloads and converts an image from a URL to a PIL Image object. + """ + response = requests.get(url, verify=False, timeout=20) + image = Image.open(BytesIO(response.content)) + return image.convert("RGB") + + def collate_fn(example, image_column="image_url", text_column="caption"): + """ + Preprocesses an example by loading and transforming image and text data. + Checks if the text data in the example is valid by calling the `check_text_data` function. + Downloads the image specified by the URL in the image_column by calling the `get_pil_from_url` function. + If there is any error during the download process, returns None. + Returns the preprocessed inputs with transformed image and text data. + """ + assert len(example) == 1 + example = example[0] + + if not check_text_data(example[text_column]): + raise ValueError("Text data is not valid") + + url = example[image_column] + try: + image = get_pil_from_url(url) + except Exception: + return None + + inputs = processor(text=example[text_column], images=[image], return_tensors="pt", padding=True) + if inputs['input_ids'].shape[1] > max_length: + return None + return inputs + +.. code:: ipython3 + + import torch + from datasets import load_dataset + + def prepare_calibration_data(dataloader, init_steps): + """ + This function prepares calibration data from a dataloader for a specified number of initialization steps. + It iterates over the dataloader, fetching batches and storing the relevant data. + """ + data = [] + print(f"Fetching {init_steps} for the initialization...") + counter = 0 + for batch in dataloader: + if counter == init_steps: + break + if batch: + counter += 1 + with torch.no_grad(): + data.append( + { + "pixel_values": batch["pixel_values"].to("cpu"), + "input_ids": batch["input_ids"].to("cpu"), + "attention_mask": batch["attention_mask"].to("cpu") + } + ) + return data + + + def prepare_dataset(opt_init_steps=300, max_train_samples=1000): + """ + Prepares a vision-text dataset for quantization. + """ + dataset = load_dataset("conceptual_captions", streaming=True) + train_dataset = dataset["train"].shuffle(seed=42, buffer_size=max_train_samples) + dataloader = torch.utils.data.DataLoader(train_dataset, collate_fn=collate_fn, batch_size=1) + calibration_data = prepare_calibration_data(dataloader, opt_init_steps) + return calibration_data + +Create a quantized model from the pre-trained ``FP16`` model. + + **NOTE**: Quantization is time and memory consuming operation. + Running quantization code below may take a long time. + +.. code:: ipython3 + + import logging + import nncf + from openvino.runtime import Core, serialize + + core = Core() + + nncf.set_log_level(logging.ERROR) + + int8_model_path = 'clip-vit-base-patch16_int8.xml' + calibration_data = prepare_dataset() + ov_model = core.read_model(fp16_model_path) + + +.. parsed-literal:: + + INFO:nncf:NNCF initialized successfully. Supported frameworks detected: torch, onnx, openvino + + + +.. parsed-literal:: + + Downloading builder script: 0%| | 0.00/6.69k [00:00`__ +in the NNCF repository for more information. + +Run quantized OpenVINO model `⇑ <#top>`__ +############################################################################################################################### + + +The steps for making predictions with the quantized OpenVINO CLIP model +are similar to the PyTorch model. Let us check the model result using +the same input data from the `1st +notebook <228-clip-zero-shot-image-classification.ipynb>`__. + +.. code:: ipython3 + + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=3, options=('CPU', 'GPU.0', 'GPU.1', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + import numpy as np + from scipy.special import softmax + from openvino.runtime import compile_model + from visualize import visualize_result + + image = Image.open('../data/image/coco.jpg') + input_labels = ['cat', 'dog', 'wolf', 'tiger', 'man', 'horse', 'frog', 'tree', 'house', 'computer'] + text_descriptions = [f"This is a photo of a {label}" for label in input_labels] + + inputs = processor(text=text_descriptions, images=[image], return_tensors="pt", padding=True) + compiled_model = compile_model(int8_model_path) + logits_per_image_out = compiled_model.output(0) + ov_logits_per_image = compiled_model(dict(inputs))[logits_per_image_out] + probs = softmax(ov_logits_per_image, axis=1) + visualize_result(image, input_labels, probs[0]) + + + +.. image:: 228-clip-zero-shot-quantize-with-output_files/228-clip-zero-shot-quantize-with-output_16_0.png + + +Compare File Size `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + + +.. code:: ipython3 + + from pathlib import Path + + fp16_ir_model_size = Path(fp16_model_path).with_suffix(".bin").stat().st_size / 1024 / 1024 + quantized_model_size = Path(int8_model_path).with_suffix(".bin").stat().st_size / 1024 / 1024 + print(f"FP16 IR model size: {fp16_ir_model_size:.2f} MB") + print(f"INT8 model size: {quantized_model_size:.2f} MB") + print(f"Model compression rate: {fp16_ir_model_size / quantized_model_size:.3f}") + + +.. parsed-literal:: + + FP16 IR model size: 285.38 MB + INT8 model size: 168.14 MB + Model compression rate: 1.697 + + +Compare inference time of the FP16 IR and quantized models +`⇑ <#top>`__ To measure the inference performance of the ``FP16`` and +``INT8`` models, we use median inference time on calibration dataset. So +we can approximately estimate the speed up of the dynamic quantized +models. + + **NOTE**: For the most accurate performance estimation, it is + recommended to run ``benchmark_app`` in a terminal/command prompt + after closing other applications with static shapes. + +.. code:: ipython3 + + import time + from openvino.runtime import compile_model + + def calculate_inference_time(model_path, calibration_data): + model = compile_model(model_path) + output_layer = model.output(0) + inference_time = [] + for batch in calibration_data: + start = time.perf_counter() + _ = model(batch)[output_layer] + end = time.perf_counter() + delta = end - start + inference_time.append(delta) + return np.median(inference_time) + +.. code:: ipython3 + + fp16_latency = calculate_inference_time(fp16_model_path, calibration_data) + int8_latency = calculate_inference_time(int8_model_path, calibration_data) + print(f"Performance speed up: {fp16_latency / int8_latency:.3f}") + + +.. parsed-literal:: + + Performance speed up: 2.092 + diff --git a/docs/notebooks/228-clip-zero-shot-quantize-with-output_files/228-clip-zero-shot-quantize-with-output_16_0.png b/docs/notebooks/228-clip-zero-shot-quantize-with-output_files/228-clip-zero-shot-quantize-with-output_16_0.png new file mode 100644 index 00000000000..eb2f7d94e98 --- /dev/null +++ b/docs/notebooks/228-clip-zero-shot-quantize-with-output_files/228-clip-zero-shot-quantize-with-output_16_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:f59ef4796e26565ffd05fe4ad9681c3188952ba6bb297319aeb75e29dd86fafa +size 464309 diff --git a/docs/notebooks/228-clip-zero-shot-quantize-with-output_files/index.html b/docs/notebooks/228-clip-zero-shot-quantize-with-output_files/index.html new file mode 100644 index 00000000000..5f79389b92c --- /dev/null +++ b/docs/notebooks/228-clip-zero-shot-quantize-with-output_files/index.html @@ -0,0 +1,7 @@ + +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/228-clip-zero-shot-quantize-with-output_files/ + +

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/228-clip-zero-shot-quantize-with-output_files/


../
+228-clip-zero-shot-quantize-with-output_16_0.png   16-Aug-2023 01:31              464309
+

+ diff --git a/docs/notebooks/229-distilbert-sequence-classification-with-output.rst b/docs/notebooks/229-distilbert-sequence-classification-with-output.rst index 40d06d9043e..514d49925a5 100644 --- a/docs/notebooks/229-distilbert-sequence-classification-with-output.rst +++ b/docs/notebooks/229-distilbert-sequence-classification-with-output.rst @@ -1,14 +1,31 @@ Sentiment Analysis with OpenVINO™ ================================= +.. _top: + **Sentiment analysis** is the use of natural language processing, text analysis, computational linguistics, and biometrics to systematically identify, extract, quantify, and study affective states and subjective information. This notebook demonstrates how to convert and run a sequence classification model using OpenVINO. -Imports -------- +**Table of contents**: + +- `Imports <#imports>`__ +- `Initializing the Model <#initializing-the-model>`__ +- `Initializing the Tokenizer <#initializing-the-tokenizer>`__ +- `Convert Model to OpenVINO Intermediate Representation format <#convert-model-to-openvino-intermediate-representation-format>`__ + + - `Select inference device <#select-inference-device>`__ + +- `Inference <#inference>`__ + + - `For a single input sentence <#for single -a- -input-sentence>`__ + - `Read from a text file <#read-from-a-text-file>`__ + +Imports `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -20,11 +37,11 @@ Imports from openvino.tools import mo from openvino.runtime import PartialShape, Type, serialize, Core -Initializing the Model ----------------------- +Initializing the Model `⇑ <#top>`__ +############################################################################################################################### -We will use the transformer-based -`distilbert-base-uncased-finetuned-sst-2-english `__ +We will use the transformer-based +`DistilBERT base uncased finetuned SST-2 `__ model from Hugging Face. .. code:: ipython3 @@ -34,8 +51,9 @@ model from Hugging Face. pretrained_model_name_or_path=checkpoint ) -Initializing the Tokenizer --------------------------- +Initializing the Tokenizer `⇑ <#top>`__ +############################################################################################################################### + Text Preprocessing cleans the text-based input data so it can be fed into the model. @@ -46,7 +64,7 @@ tokens or IDs to the words, so they are represented in a vector space where similar words have similar vectors. This helps the model understand the context of a sentence. Here, we will use `AutoTokenizer `__ -- a pre-trained tokenizer from Hugging Face: . +- a pre-trained tokenizer from Hugging Face: .. code:: ipython3 @@ -54,15 +72,13 @@ understand the context of a sentence. Here, we will use pretrained_model_name_or_path=checkpoint ) -Convert Model to OpenVINO Intermediate Representation format ------------------------------------------------------------- +Convert Model to OpenVINO Intermediate Representation format. `⇑ <#top>`__ +############################################################################################################################### -`Model -Optimizer `__ -is a cross-platform command-line tool that facilitates the transition -between training and deployment environments, performs static model -analysis, and adjusts deep learning models for optimal execution on -end-point target devices. +`Model conversion API `__ +facilitates the transition between training and deployment environments, +performs static model analysis, and adjusts deep learning models for +optimal execution on end-point target devices. .. code:: ipython3 @@ -75,12 +91,11 @@ end-point target devices. .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/distilbert/modeling_distilbert.py:223: TracerWarning: torch.tensor results are registered as constants in the trace. You can safely ignore this warning if you use this function to create tensors out of constant variables that would be the same every time you call this function. In any other case, this might cause the trace to be incorrect. + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/distilbert/modeling_distilbert.py:223: TracerWarning: torch.tensor results are registered as constants in the trace. You can safely ignore this warning if you use this function to create tensors out of constant variables that would be the same every time you call this function. In any other case, this might cause the trace to be incorrect. mask, torch.tensor(torch.finfo(scores.dtype).min) -OpenVINO™ Runtime uses the `Infer -Request `__ +OpenVINO™ Runtime uses the `Infer Request `__ mechanism which enables running models on different devices in asynchronous or synchronous manners. The model graph is sent as an argument to the OpenVINO API and an inference request is created. The @@ -91,9 +106,40 @@ documentation. `__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + warnings.filterwarnings("ignore") + compiled_model = core.compile_model(ov_model, device.value) infer_request = compiled_model.create_infer_request() .. code:: ipython3 @@ -109,8 +155,9 @@ documentation. `__ +############################################################################################################################### + .. code:: ipython3 @@ -135,8 +182,9 @@ Inference probability = np.argmax(softmax(i)) return label[probability] -For a single input sentence -~~~~~~~~~~~~~~~~~~~~~~~~~~~ +For a single input sentence `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -155,8 +203,9 @@ For a single input sentence Total Time: 0.04 seconds -Read from a text file -~~~~~~~~~~~~~~~~~~~~~ +Read from a text file `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 diff --git a/docs/notebooks/230-yolov8-optimization-with-output.rst b/docs/notebooks/230-yolov8-optimization-with-output.rst index fc75b6a1fe4..e27522d664d 100644 --- a/docs/notebooks/230-yolov8-optimization-with-output.rst +++ b/docs/notebooks/230-yolov8-optimization-with-output.rst @@ -1,6 +1,8 @@ Convert and Optimize YOLOv8 with OpenVINO™ ========================================== +.. _top: + The YOLOv8 algorithm developed by Ultralytics is a cutting-edge, state-of-the-art (SOTA) model that is designed to be fast, accurate, and easy to use, making it an excellent choice for a wide range of object @@ -28,17 +30,67 @@ for object detection and instance segmentation scenarios. The tutorial consists of the following steps: -- Prepare the PyTorch model. -- Download and prepare a dataset. -- Validate the original model. -- Convert the PyTorch model to OpenVINO IR. -- Validate the converted model. -- Prepare and run optimization pipeline. -- Compare performance of the FP32 and quantized models. -- Compare accuracy of the FP32 and quantized models. +- Prepare the PyTorch model. +- Download and prepare a dataset. +- Validate the original model. +- Convert the PyTorch model to OpenVINO IR. +- Validate the converted model. +- Prepare and run optimization pipeline. +- Compare performance of the FP32 and quantized models. +- Compare accuracy of the FP32 and quantized models. + +**Table of contents**: + +- `Get Pytorch model <#get-pytorch-model>`__ +- `Prerequisites <#prerequisites>`__ +- `Instantiate model <#instantiate-model>`__ + + - `Object detection <#object-detection>`__ + - `Instance Segmentation: <#instance-segmentation>`__ + - `Convert model to OpenVINO IR <#convert-model-to-openvino-ir>`__ + - `Verify model inference <#verify-model-inference>`__ + - `Preprocessing <#preprocessing>`__ + - `Postprocessing <#postprocessing>`__ + - `Select inference device <#select-inference-device>`__ + - `Test on single image <#test-on-single-image>`__ + - `Check model accuracy on the dataset <#check-model-accuracy-on-the-dataset>`__ + + - `Download the validation dataset <#download-the-validation-dataset>`__ + - `Define validation function <#define-validation-function>`__ + - `Configure Validator helper and create DataLoader <#configure-validator-helper-and-create-dataloader>`__ + + - `Optimize model using NNCF Post-training Quantization API <#optimize-model-using-nncf-post-training-quantization-api>`__ + - `Validate Quantized model inference <#validate-quantized-model-inference>`__ + + - `Object detection: <#object-detection>`__ + - `Instance segmentation: <#instance-segmentation>`__ + + - `Compare Performance of the Original and Quantized Models <#compare-performance-of-the-original-and-quantized-models>`__ + + - `Compare performance object detection models <#compare-performance-object-detection-models>`__ + - `Instance segmentation <#instance-segmentation>`__ + + - `Validate quantized model accuracy <#validate-quantized-model-accuracy>`__ + - `Object detection <#object-detection>`__ + - `Instance segmentation <#instance-segmentation>`__ + +- `Next steps <#next-steps>`__ +- `Async inference pipeline <#async-inference-pipeline>`__ +- `Integration preprocessing to model <#integration-preprocessing-to-model>`__ + + - `Initialize PrePostProcessing API <#initialize-prepostprocessing-api>`__ + - `Define input data format <#define-input-data-format>`__ + - `Describe preprocessing steps <#describe-preprocessing-steps>`__ + - `Integrating Steps into a Model <#integrating-steps-into-a-model>`__ + +- `Live demo <#live-demo>`__ +- `Run <#run>`__ + + - `Run Live Object Detection and Segmentation <#run-live-object-detection-and-segmentation>`__ + +Get Pytorch model `⇑ <#top>`__ +############################################################################################################################### -Get Pytorch model ------------------ Generally, PyTorch models represent an instance of the `torch.nn.Module `__ @@ -50,22 +102,25 @@ also applicable to other YOLOv8 models. Typical steps to obtain a pre-trained model: 1. Create an instance of a model class. -2. Load a checkpoint state dict, which contains the pre-trained model weights. -3. Turn the model to evaluation for switching some operations to inference mode. +2. Load a checkpoint state dict, which contains the pre-trained model + weights. +3. Turn the model to evaluation for switching some operations to + inference mode. In this case, the creators of the model provide an API that enables converting the YOLOv8 model to ONNX and then to OpenVINO IR. Therefore, we do not need to do these steps manually. -Prerequisites -^^^^^^^^^^^^^ +Prerequisites `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + Install necessary packages. .. code:: ipython3 - !pip install -q 'openvino-dev>=2023.0.0' 'nncf>=2.5.0' - !pip install -q 'ultralytics==8.0.43' onnx + !pip install -q "openvino-dev>=2023.0.0" "nncf>=2.5.0" + !pip install -q "ultralytics==8.0.43" onnx Import required utility functions. The lower cell will download the ``notebook_utils`` Python module from GitHub. @@ -163,12 +218,13 @@ Define utility functions for drawing results .. parsed-literal:: - PosixPath('/home/idavidyu/openvino_notebooks/notebooks/230-yolov8-optimization/data/coco_bike.jpg') + PosixPath('/home/ea/work/openvino_notebooks/notebooks/230-yolov8-optimization/data/coco_bike.jpg') -Instantiate model ------------------ +Instantiate model `⇑ <#top>`__ +############################################################################################################################### + There are several models available in the original repository, targeted for different tasks. For loading the model, required to specify a path @@ -189,8 +245,9 @@ Let us consider the examples: models_dir = Path('./models') models_dir.mkdir(exist_ok=True) -Object detection -~~~~~~~~~~~~~~~~ +Object detection `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -207,21 +264,22 @@ Object detection .. parsed-literal:: - Ultralytics YOLOv8.0.43 🚀 Python-3.10.6 torch-2.0.1+cu117 CUDA:0 (NVIDIA GeForce RTX 3090, 24260MiB) + Ultralytics YOLOv8.0.43 🚀 Python-3.8.10 torch-1.13.1+cpu CPU YOLOv8n summary (fused): 168 layers, 3151904 parameters, 0 gradients, 8.7 GFLOPs - image 1/1 /home/idavidyu/openvino_notebooks/notebooks/230-yolov8-optimization/data/coco_bike.jpg: 480x640 2 bicycles, 2 cars, 1 dog, 61.5ms - Speed: 1.4ms preprocess, 61.5ms inference, 1.2ms postprocess per image at shape (1, 3, 640, 640) + image 1/1 /home/ea/work/openvino_notebooks/notebooks/230-yolov8-optimization/data/coco_bike.jpg: 480x640 2 bicycles, 2 cars, 1 dog, 43.6ms + Speed: 0.5ms preprocess, 43.6ms inference, 1.0ms postprocess per image at shape (1, 3, 640, 640) -.. image:: 230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_12_1.png +.. image:: 230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_13_1.png -Instance Segmentation: -~~~~~~~~~~~~~~~~~~~~~~ +Instance Segmentation: `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -234,23 +292,22 @@ Instance Segmentation: .. parsed-literal:: - Ultralytics YOLOv8.0.43 🚀 Python-3.10.6 torch-2.0.1+cu117 CUDA:0 (NVIDIA GeForce RTX 3090, 24260MiB) + Ultralytics YOLOv8.0.43 🚀 Python-3.8.10 torch-1.13.1+cpu CPU YOLOv8n-seg summary (fused): 195 layers, 3404320 parameters, 0 gradients, 12.6 GFLOPs - image 1/1 /home/idavidyu/openvino_notebooks/notebooks/230-yolov8-optimization/data/coco_bike.jpg: 480x640 1 bicycle, 2 cars, 1 dog, 19.3ms - Speed: 0.3ms preprocess, 19.3ms inference, 1.4ms postprocess per image at shape (1, 3, 640, 640) - /home/idavidyu/.virtualenvs/test/lib/python3.10/site-packages/torchvision/transforms/functional.py:1603: UserWarning: The default value of the antialias parameter of all the resizing transforms (Resize(), RandomResizedCrop(), etc.) will change from None to True in v0.17, in order to be consistent across the PIL and Tensor backends. To suppress this warning, directly pass antialias=True (recommended, future default), antialias=None (current default, which means False for Tensors and True for PIL), or antialias=False (only works on Tensors - PIL will still use antialiasing). This also applies if you are using the inference transforms from the models weights: update the call to weights.transforms(antialias=True). - warnings.warn( + image 1/1 /home/ea/work/openvino_notebooks/notebooks/230-yolov8-optimization/data/coco_bike.jpg: 480x640 1 bicycle, 2 cars, 1 dog, 43.2ms + Speed: 0.5ms preprocess, 43.2ms inference, 1.6ms postprocess per image at shape (1, 3, 640, 640) -.. image:: 230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_14_1.png +.. image:: 230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_15_1.png -Convert model to OpenVINO IR -~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Convert model to OpenVINO IR `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + YOLOv8 provides API for convenient model exporting to different formats including OpenVINO IR. ``model.export`` is responsible for model @@ -271,8 +328,9 @@ preserve dynamic shapes in the model. if not seg_model_path.exists(): seg_model.export(format="openvino", dynamic=True, half=False) -Verify model inference -~~~~~~~~~~~~~~~~~~~~~~ +Verify model inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + To test model work, we create inference pipeline similar to ``model.predict`` method. The pipeline consists of preprocessing step, @@ -281,16 +339,17 @@ The main difference in models for object detection and instance segmentation is postprocessing part. Input specification and preprocessing are common for both cases. -Preprocessing -~~~~~~~~~~~~~ +Preprocessing `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Model input is a tensor with the ``[-1, 3, -1, -1]`` shape in the ``N, C, H, W`` format, where -* ``N`` - number of images in batch (batch size) -* ``C`` - image channels -* ``H`` - image height -* ``W`` - image width +- ``N`` - number of images in batch (batch size) +- ``C`` - image channels +- ``H`` - image height +- ``W`` - image width The model expects images in RGB channels format and normalized in [0, 1] range. Although the model supports dynamic input shape with preserving @@ -398,8 +457,9 @@ To keep a specific shape, preprocessing automatically enables padding. input_tensor = np.expand_dims(input_tensor, 0) return input_tensor -Postprocessing -~~~~~~~~~~~~~~ +Postprocessing `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The model output contains detection boxes candidates, it is a tensor with the ``[-1,84,-1]`` shape in the ``B,84,N`` format, where: @@ -424,14 +484,18 @@ contains proto mask candidates for instance segmentation. It should be decoded by using box coordinates. It is a tensor with the ``[-1 32, -1, -1]`` shape in the ``B,C H,W`` format, where: -- ``B`` - batch size -- ``C`` - number of candidates -- ``H`` - mask height -- ``W`` - mask width - +- ``B`` - batch size +- ``C`` - number of candidates +- ``H`` - mask height +- ``W`` - mask width .. code:: ipython3 + try: + scale_segments = ops.scale_segments + except AttributeError: + scale_segments = ops.scale_coords + def postprocess( pred_boxes:np.ndarray, input_hw:Tuple[int, int], @@ -483,16 +547,48 @@ decoded by using box coordinates. It is a tensor with the if retina_mask: pred[:, :4] = ops.scale_boxes(input_hw, pred[:, :4], shape).round() masks = ops.process_mask_native(proto[i], pred[:, 6:], pred[:, :4], shape[:2]) # HWC - segments = [ops.scale_segments(input_hw, x, shape, normalize=False) for x in ops.masks2segments(masks)] + segments = [scale_segments(input_hw, x, shape, normalize=False) for x in ops.masks2segments(masks)] else: masks = ops.process_mask(proto[i], pred[:, 6:], pred[:, :4], input_hw, upsample=True) pred[:, :4] = ops.scale_boxes(input_hw, pred[:, :4], shape).round() - segments = [ops.scale_segments(input_hw, x, shape, normalize=False) for x in ops.masks2segments(masks)] + segments = [scale_segments(input_hw, x, shape, normalize=False) for x in ops.masks2segments(masks)] results.append({"det": pred[:, :6].numpy(), "segment": segments}) return results -Test on single image -~~~~~~~~~~~~~~~~~~~~ +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + from openvino.runtime import Core + + core = Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=2, options=('CPU', 'GPU', 'AUTO'), value='AUTO') + + + +Test on single image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Now, once we have defined preprocessing and postprocessing steps, we are ready to check model prediction. @@ -505,10 +601,9 @@ First, object detection: core = Core() det_ov_model = core.read_model(det_model_path) - device = "CPU" # "GPU" - if device != "CPU": + if device.value != "CPU": det_ov_model.reshape({0: [1, 3, 640, 640]}) - det_compiled_model = core.compile_model(det_ov_model, device) + det_compiled_model = core.compile_model(det_ov_model, device.value) def detect(image:np.ndarray, model:Model): @@ -542,7 +637,7 @@ First, object detection: -.. image:: 230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_24_0.png +.. image:: 230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_27_0.png @@ -551,10 +646,9 @@ Then, instance segmentation: .. code:: ipython3 seg_ov_model = core.read_model(seg_model_path) - device = "CPU" # GPU - if device != "CPU": + if device.value != "CPU": seg_ov_model.reshape({0: [1, 3, 640, 640]}) - seg_compiled_model = core.compile_model(seg_ov_model, device) + seg_compiled_model = core.compile_model(seg_ov_model, device.value) input_image = np.array(Image.open(IMAGE_PATH)) @@ -567,21 +661,23 @@ Then, instance segmentation: -.. image:: 230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_26_0.png +.. image:: 230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_29_0.png Great! The result is the same, as produced by original models. -Check model accuracy on the dataset -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Check model accuracy on the dataset `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + For comparing the optimized model result with the original, it is good to know some measurable results in terms of model accuracy on the validation dataset. -Download the validation dataset -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +Download the validation dataset `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + YOLOv8 is pre-trained on the COCO dataset, so to evaluate the model accuracy we need to download it. According to the instructions provided @@ -599,7 +695,7 @@ evaluation function. DATA_URL = "http://images.cocodataset.org/zips/val2017.zip" LABELS_URL = "https://github.com/ultralytics/yolov5/releases/download/v1.0/coco2017labels-segments.zip" - CFG_URL = "https://raw.githubusercontent.com/ultralytics/ultralytics/main/ultralytics/datasets/coco.yaml" + CFG_URL = "https://raw.githubusercontent.com/ultralytics/ultralytics/main/ultralytics/cfg/datasets/coco.yaml" OUT_DIR = Path('./datasets') @@ -630,8 +726,9 @@ evaluation function. datasets/coco.yaml: 0%| | 0.00/1.25k [00:00`__ +------------------------------------------------------------------------------------------------------------------------------- + .. code:: ipython3 @@ -698,8 +795,9 @@ Define validation function pf = '%20s' + '%12i' * 2 + '%12.3g' * 4 # print format print(pf % ('all', total_images, total_objects, s_mp, s_mr, s_map50, s_mean_ap)) -Configure Validator helper and create DataLoader -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +Configure Validator helper and create DataLoader `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + The original model repository uses a ``Validator`` wrapper, which represents the accuracy validation pipeline. It creates dataloader and @@ -765,14 +863,13 @@ validator class instance. After definition test function and validator creation, we are ready for -getting accuracy metrics. - -.. note:: - - Model evaluation is time consuming process and can take several minutes, depending on the hardware. For reducing calculation time, we define ``num_samples`` parameter with evaluation subset size, but in this case, accuracy can be noncomparable with originally reported by the authors of the model, due to validation subset difference. - - -*To validate the models on the full dataset set* ``NUM_TEST_SAMPLES = None``. +getting accuracy metrics >\ **Note**: Model evaluation is time consuming +process and can take several minutes, depending on the hardware. For +reducing calculation time, we define ``num_samples`` parameter with +evaluation subset size, but in this case, accuracy can be noncomparable +with originally reported by the authors of the model, due to validation +subset difference. *To validate the models on the full dataset set +``NUM_TEST_SAMPLES = None``.* .. code:: ipython3 @@ -786,7 +883,7 @@ getting accuracy metrics. .. parsed-literal:: - 0%| | 0/500 [00:00`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + `NNCF `__ provides a suite of advanced algorithms for Neural Networks inference optimization in @@ -884,7 +982,15 @@ for both models is the same, we can reuse one dataset for both models. .. parsed-literal:: - INFO:nncf:NNCF initialized successfully. Supported frameworks detected: torch, onnx, openvino + 2023-07-14 18:41:29.274964: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-07-14 18:41:29.313487: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. + 2023-07-14 18:41:29.989212: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + + +.. parsed-literal:: + + INFO:nncf:NNCF initialized successfully. Supported frameworks detected: torch, tensorflow, onnx, openvino The ``nncf.quantize`` function provides an interface for model @@ -947,8 +1053,8 @@ point precision, using the ``ignored_scope`` parameter. .. parsed-literal:: - Statistics collection: 100%|███████████████████████████████████████████████████████████████████████████| 300/300 [00:26<00:00, 11.36it/s] - Biases correction: 100%|█████████████████████████████████████████████████████████████████████████████████| 63/63 [00:02<00:00, 29.82it/s] + Statistics collection: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 300/300 [00:34<00:00, 8.79it/s] + Biases correction: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 63/63 [00:02<00:00, 22.46it/s] .. code:: ipython3 @@ -991,8 +1097,8 @@ point precision, using the ``ignored_scope`` parameter. .. parsed-literal:: - Statistics collection: 100%|███████████████████████████████████████████████████████████████████████████| 300/300 [00:31<00:00, 9.48it/s] - Biases correction: 100%|█████████████████████████████████████████████████████████████████████████████████| 75/75 [00:02<00:00, 30.19it/s] + Statistics collection: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 300/300 [00:40<00:00, 7.45it/s] + Biases correction: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 75/75 [00:03<00:00, 23.13it/s] .. code:: ipython3 @@ -1007,8 +1113,9 @@ point precision, using the ``ignored_scope`` parameter. Quantized segmentation model will be saved to models/yolov8n-seg_openvino_int8_model/yolov8n-seg.xml -Validate Quantized model inference -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Validate Quantized model inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + ``nncf.quantize`` returns the OpenVINO Model class instance, which is suitable for loading on a device for making predictions. ``INT8`` model @@ -1017,14 +1124,28 @@ floating point model representation. Therefore, we can reuse the same ``detect`` function defined above for getting the ``INT8`` model result on the image. -Object detection: -^^^^^^^^^^^^^^^^^ +.. code:: ipython3 + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=2, options=('CPU', 'GPU', 'AUTO'), value='AUTO') + + + +Object detection: `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + .. code:: ipython3 - if device != "CPU": - quantized_det_model.reshape({0, [1, 3, 640, 640]}) - quantized_det_compiled_model = core.compile_model(quantized_det_model, device) + if device.value != "CPU": + quantized_det_model.reshape({0: [1, 3, 640, 640]}) + quantized_det_compiled_model = core.compile_model(quantized_det_model, device.value) input_image = np.array(Image.open(IMAGE_PATH)) detections = detect(input_image, quantized_det_compiled_model)[0] image_with_boxes = draw_results(detections, input_image, label_map) @@ -1034,18 +1155,19 @@ Object detection: -.. image:: 230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_54_0.png +.. image:: 230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_59_0.png -Instance segmentation: -^^^^^^^^^^^^^^^^^^^^^^ +Instance segmentation: `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + .. code:: ipython3 - if device != "CPU": - quantized_seg_model.reshape({0, [1, 3, 640, 640]}) - quantized_seg_compiled_model = core.compile_model(quantized_seg_model, device) + if device.value != "CPU": + quantized_seg_model.reshape({0: [1, 3, 640, 640]}) + quantized_seg_compiled_model = core.compile_model(quantized_seg_model, device.value) input_image = np.array(Image.open(IMAGE_PATH)) detections = detect(input_image, quantized_seg_compiled_model)[0] image_with_masks = draw_results(detections, input_image, label_map) @@ -1055,12 +1177,12 @@ Instance segmentation: -.. image:: 230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_56_0.png +.. image:: 230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_61_0.png -Compare Performance of the Original and Quantized Models -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Compare Performance of the Original and Quantized Models `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ Finally, use the OpenVINO `Benchmark Tool `__ @@ -1076,13 +1198,27 @@ models. ``benchmark_app --help`` to see an overview of all command-line options. -Compare performance object detection models -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +Compare performance object detection models `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + + +.. code:: ipython3 + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=2, options=('CPU', 'GPU', 'AUTO'), value='AUTO') + + .. code:: ipython3 # Inference FP32 model (OpenVINO IR) - !benchmark_app -m $det_model_path -d $device -api async -shape "[1,3,640,640]" + !benchmark_app -m $det_model_path -d $device.value -api async -shape "[1,3,640,640]" .. parsed-literal:: @@ -1090,19 +1226,20 @@ Compare performance object detection models [Step 1/11] Parsing and validating input arguments [ INFO ] Parsing input parameters [Step 2/11] Loading OpenVINO Runtime + [ WARNING ] Default duration 120 seconds is used for unknown device AUTO [ INFO ] OpenVINO: - [ INFO ] Build ................................. 2023.0.0-10926-b4452d56304-releases/2023/0 + [ INFO ] Build ................................. 2023.0.1-11005-fa1c41994f3-releases/2023/0 [ INFO ] [ INFO ] Device info: - [ INFO ] CPU - [ INFO ] Build ................................. 2023.0.0-10926-b4452d56304-releases/2023/0 + [ INFO ] AUTO + [ INFO ] Build ................................. 2023.0.1-11005-fa1c41994f3-releases/2023/0 [ INFO ] [ INFO ] [Step 3/11] Setting device configuration - [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.THROUGHPUT. + [ WARNING ] Performance hint was not explicitly specified in command line. Device(AUTO) performance hint will be set to PerformanceMode.THROUGHPUT. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 17.85 ms + [ INFO ] Read model took 16.88 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] images (node: images) : f32 / [...] / [?,3,?,?] @@ -1111,52 +1248,59 @@ Compare performance object detection models [Step 5/11] Resizing model to match image sizes and given batch [ INFO ] Model batch size: 1 [ INFO ] Reshaping model: 'images': [1,3,640,640] - [ INFO ] Reshape model took 9.98 ms + [ INFO ] Reshape model took 11.45 ms [Step 6/11] Configuring input of the model [ INFO ] Model inputs: [ INFO ] images (node: images) : u8 / [N,C,H,W] / [1,3,640,640] [ INFO ] Model outputs: [ INFO ] output0 (node: output0) : f32 / [...] / [1,84,8400] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 213.37 ms + [ INFO ] Compile model took 410.99 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: + [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT [ INFO ] NETWORK_NAME: torch_jit [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 12 - [ INFO ] NUM_STREAMS: 12 - [ INFO ] AFFINITY: Affinity.CORE - [ INFO ] INFERENCE_NUM_THREADS: 36 - [ INFO ] PERF_COUNT: False - [ INFO ] INFERENCE_PRECISION_HINT: - [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT - [ INFO ] EXECUTION_MODE_HINT: ExecutionMode.PERFORMANCE - [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 - [ INFO ] ENABLE_CPU_PINNING: True - [ INFO ] SCHEDULING_CORE_TYPE: SchedulingCoreType.ANY_CORE - [ INFO ] ENABLE_HYPER_THREADING: True + [ INFO ] MODEL_PRIORITY: Priority.MEDIUM + [ INFO ] MULTI_DEVICE_PRIORITIES: CPU + [ INFO ] CPU: + [ INFO ] CPU_BIND_THREAD: YES + [ INFO ] CPU_THREADS_NUM: 0 + [ INFO ] CPU_THROUGHPUT_STREAMS: 12 + [ INFO ] DEVICE_ID: + [ INFO ] DUMP_EXEC_GRAPH_AS_DOT: + [ INFO ] DYN_BATCH_ENABLED: NO + [ INFO ] DYN_BATCH_LIMIT: 0 + [ INFO ] ENFORCE_BF16: NO + [ INFO ] EXCLUSIVE_ASYNC_REQUESTS: NO + [ INFO ] NETWORK_NAME: torch_jit + [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 12 + [ INFO ] PERFORMANCE_HINT: THROUGHPUT + [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 + [ INFO ] PERF_COUNT: NO [ INFO ] EXECUTION_DEVICES: ['CPU'] [Step 9/11] Creating infer requests and preparing input tensors [ WARNING ] No input files were given for input 'images'!. This input will be filled with random values! [ INFO ] Fill input 'images' with random values - [Step 10/11] Measuring performance (Start inference asynchronously, 12 inference requests, limits: 60000 ms duration) + [Step 10/11] Measuring performance (Start inference asynchronously, 12 inference requests, limits: 120000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 32.22 ms + [ INFO ] First inference took 30.16 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 11544 iterations - [ INFO ] Duration: 60045.35 ms + [ INFO ] Count: 19752 iterations + [ INFO ] Duration: 120070.55 ms [ INFO ] Latency: - [ INFO ] Median: 62.15 ms - [ INFO ] Average: 62.22 ms - [ INFO ] Min: 40.94 ms - [ INFO ] Max: 83.81 ms - [ INFO ] Throughput: 192.25 FPS + [ INFO ] Median: 71.27 ms + [ INFO ] Average: 72.76 ms + [ INFO ] Min: 47.53 ms + [ INFO ] Max: 164.37 ms + [ INFO ] Throughput: 164.50 FPS .. code:: ipython3 # Inference INT8 model (OpenVINO IR) - !benchmark_app -m $int8_model_det_path -d $device -api async -shape "[1,3,640,640]" -t 15 + !benchmark_app -m $int8_model_det_path -d $device.value -api async -shape "[1,3,640,640]" -t 15 .. parsed-literal:: @@ -1165,18 +1309,18 @@ Compare performance object detection models [ INFO ] Parsing input parameters [Step 2/11] Loading OpenVINO Runtime [ INFO ] OpenVINO: - [ INFO ] Build ................................. 2023.0.0-10926-b4452d56304-releases/2023/0 + [ INFO ] Build ................................. 2023.0.1-11005-fa1c41994f3-releases/2023/0 [ INFO ] [ INFO ] Device info: - [ INFO ] CPU - [ INFO ] Build ................................. 2023.0.0-10926-b4452d56304-releases/2023/0 + [ INFO ] AUTO + [ INFO ] Build ................................. 2023.0.1-11005-fa1c41994f3-releases/2023/0 [ INFO ] [ INFO ] [Step 3/11] Setting device configuration - [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.THROUGHPUT. + [ WARNING ] Performance hint was not explicitly specified in command line. Device(AUTO) performance hint will be set to PerformanceMode.THROUGHPUT. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 25.95 ms + [ INFO ] Read model took 27.47 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] images (node: images) : f32 / [...] / [1,3,?,?] @@ -1185,54 +1329,62 @@ Compare performance object detection models [Step 5/11] Resizing model to match image sizes and given batch [ INFO ] Model batch size: 1 [ INFO ] Reshaping model: 'images': [1,3,640,640] - [ INFO ] Reshape model took 16.79 ms + [ INFO ] Reshape model took 14.87 ms [Step 6/11] Configuring input of the model [ INFO ] Model inputs: [ INFO ] images (node: images) : u8 / [N,C,H,W] / [1,3,640,640] [ INFO ] Model outputs: [ INFO ] output0 (node: output0) : f32 / [...] / [1,84,8400] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 417.84 ms + [ INFO ] Compile model took 681.89 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: + [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT [ INFO ] NETWORK_NAME: torch_jit [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 18 - [ INFO ] NUM_STREAMS: 18 - [ INFO ] AFFINITY: Affinity.CORE - [ INFO ] INFERENCE_NUM_THREADS: 36 - [ INFO ] PERF_COUNT: False - [ INFO ] INFERENCE_PRECISION_HINT: - [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT - [ INFO ] EXECUTION_MODE_HINT: ExecutionMode.PERFORMANCE - [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 - [ INFO ] ENABLE_CPU_PINNING: True - [ INFO ] SCHEDULING_CORE_TYPE: SchedulingCoreType.ANY_CORE - [ INFO ] ENABLE_HYPER_THREADING: True + [ INFO ] MODEL_PRIORITY: Priority.MEDIUM + [ INFO ] MULTI_DEVICE_PRIORITIES: CPU + [ INFO ] CPU: + [ INFO ] CPU_BIND_THREAD: YES + [ INFO ] CPU_THREADS_NUM: 0 + [ INFO ] CPU_THROUGHPUT_STREAMS: 18 + [ INFO ] DEVICE_ID: + [ INFO ] DUMP_EXEC_GRAPH_AS_DOT: + [ INFO ] DYN_BATCH_ENABLED: NO + [ INFO ] DYN_BATCH_LIMIT: 0 + [ INFO ] ENFORCE_BF16: NO + [ INFO ] EXCLUSIVE_ASYNC_REQUESTS: NO + [ INFO ] NETWORK_NAME: torch_jit + [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 18 + [ INFO ] PERFORMANCE_HINT: THROUGHPUT + [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 + [ INFO ] PERF_COUNT: NO [ INFO ] EXECUTION_DEVICES: ['CPU'] [Step 9/11] Creating infer requests and preparing input tensors [ WARNING ] No input files were given for input 'images'!. This input will be filled with random values! [ INFO ] Fill input 'images' with random values [Step 10/11] Measuring performance (Start inference asynchronously, 18 inference requests, limits: 15000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 24.13 ms + [ INFO ] First inference took 20.61 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 7506 iterations - [ INFO ] Duration: 15053.56 ms + [ INFO ] Count: 6282 iterations + [ INFO ] Duration: 15065.20 ms [ INFO ] Latency: - [ INFO ] Median: 35.50 ms - [ INFO ] Average: 35.90 ms - [ INFO ] Min: 24.17 ms - [ INFO ] Max: 51.86 ms - [ INFO ] Throughput: 498.62 FPS + [ INFO ] Median: 41.71 ms + [ INFO ] Average: 42.98 ms + [ INFO ] Min: 25.38 ms + [ INFO ] Max: 118.34 ms + [ INFO ] Throughput: 416.99 FPS -Instance segmentation -^^^^^^^^^^^^^^^^^^^^^ +Instance segmentation `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + .. code:: ipython3 - !benchmark_app -m $seg_model_path -d $device -api async -shape "[1,3,640,640]" -t 15 + !benchmark_app -m $seg_model_path -d $device.value -api async -shape "[1,3,640,640]" -t 15 .. parsed-literal:: @@ -1241,18 +1393,18 @@ Instance segmentation [ INFO ] Parsing input parameters [Step 2/11] Loading OpenVINO Runtime [ INFO ] OpenVINO: - [ INFO ] Build ................................. 2023.0.0-10926-b4452d56304-releases/2023/0 + [ INFO ] Build ................................. 2023.0.1-11005-fa1c41994f3-releases/2023/0 [ INFO ] [ INFO ] Device info: - [ INFO ] CPU - [ INFO ] Build ................................. 2023.0.0-10926-b4452d56304-releases/2023/0 + [ INFO ] AUTO + [ INFO ] Build ................................. 2023.0.1-11005-fa1c41994f3-releases/2023/0 [ INFO ] [ INFO ] [Step 3/11] Setting device configuration - [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.THROUGHPUT. + [ WARNING ] Performance hint was not explicitly specified in command line. Device(AUTO) performance hint will be set to PerformanceMode.THROUGHPUT. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 19.62 ms + [ INFO ] Read model took 18.86 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] images (node: images) : f32 / [...] / [?,3,?,?] @@ -1262,7 +1414,7 @@ Instance segmentation [Step 5/11] Resizing model to match image sizes and given batch [ INFO ] Model batch size: 1 [ INFO ] Reshaping model: 'images': [1,3,640,640] - [ INFO ] Reshape model took 11.07 ms + [ INFO ] Reshape model took 13.15 ms [Step 6/11] Configuring input of the model [ INFO ] Model inputs: [ INFO ] images (node: images) : u8 / [N,C,H,W] / [1,3,640,640] @@ -1270,44 +1422,51 @@ Instance segmentation [ INFO ] output0 (node: output0) : f32 / [...] / [1,116,8400] [ INFO ] output1 (node: output1) : f32 / [...] / [1,32,160,160] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 247.03 ms + [ INFO ] Compile model took 420.45 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: + [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT [ INFO ] NETWORK_NAME: torch_jit [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 12 - [ INFO ] NUM_STREAMS: 12 - [ INFO ] AFFINITY: Affinity.CORE - [ INFO ] INFERENCE_NUM_THREADS: 36 - [ INFO ] PERF_COUNT: False - [ INFO ] INFERENCE_PRECISION_HINT: - [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT - [ INFO ] EXECUTION_MODE_HINT: ExecutionMode.PERFORMANCE - [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 - [ INFO ] ENABLE_CPU_PINNING: True - [ INFO ] SCHEDULING_CORE_TYPE: SchedulingCoreType.ANY_CORE - [ INFO ] ENABLE_HYPER_THREADING: True + [ INFO ] MODEL_PRIORITY: Priority.MEDIUM + [ INFO ] MULTI_DEVICE_PRIORITIES: CPU + [ INFO ] CPU: + [ INFO ] CPU_BIND_THREAD: YES + [ INFO ] CPU_THREADS_NUM: 0 + [ INFO ] CPU_THROUGHPUT_STREAMS: 12 + [ INFO ] DEVICE_ID: + [ INFO ] DUMP_EXEC_GRAPH_AS_DOT: + [ INFO ] DYN_BATCH_ENABLED: NO + [ INFO ] DYN_BATCH_LIMIT: 0 + [ INFO ] ENFORCE_BF16: NO + [ INFO ] EXCLUSIVE_ASYNC_REQUESTS: NO + [ INFO ] NETWORK_NAME: torch_jit + [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 12 + [ INFO ] PERFORMANCE_HINT: THROUGHPUT + [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 + [ INFO ] PERF_COUNT: NO [ INFO ] EXECUTION_DEVICES: ['CPU'] [Step 9/11] Creating infer requests and preparing input tensors [ WARNING ] No input files were given for input 'images'!. This input will be filled with random values! [ INFO ] Fill input 'images' with random values [Step 10/11] Measuring performance (Start inference asynchronously, 12 inference requests, limits: 15000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 38.50 ms + [ INFO ] First inference took 39.79 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 2256 iterations - [ INFO ] Duration: 15108.66 ms + [ INFO ] Count: 1920 iterations + [ INFO ] Duration: 15131.06 ms [ INFO ] Latency: - [ INFO ] Median: 79.61 ms - [ INFO ] Average: 80.02 ms - [ INFO ] Min: 41.17 ms - [ INFO ] Max: 199.22 ms - [ INFO ] Throughput: 149.32 FPS + [ INFO ] Median: 92.12 ms + [ INFO ] Average: 94.20 ms + [ INFO ] Min: 55.80 ms + [ INFO ] Max: 154.59 ms + [ INFO ] Throughput: 126.89 FPS .. code:: ipython3 - !benchmark_app -m $int8_model_seg_path -d $device -api async -shape "[1,3,640,640]" -t 15 + !benchmark_app -m $int8_model_seg_path -d $device.value -api async -shape "[1,3,640,640]" -t 15 .. parsed-literal:: @@ -1316,18 +1475,18 @@ Instance segmentation [ INFO ] Parsing input parameters [Step 2/11] Loading OpenVINO Runtime [ INFO ] OpenVINO: - [ INFO ] Build ................................. 2023.0.0-10926-b4452d56304-releases/2023/0 + [ INFO ] Build ................................. 2023.0.1-11005-fa1c41994f3-releases/2023/0 [ INFO ] [ INFO ] Device info: - [ INFO ] CPU - [ INFO ] Build ................................. 2023.0.0-10926-b4452d56304-releases/2023/0 + [ INFO ] AUTO + [ INFO ] Build ................................. 2023.0.1-11005-fa1c41994f3-releases/2023/0 [ INFO ] [ INFO ] [Step 3/11] Setting device configuration - [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.THROUGHPUT. + [ WARNING ] Performance hint was not explicitly specified in command line. Device(AUTO) performance hint will be set to PerformanceMode.THROUGHPUT. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 29.10 ms + [ INFO ] Read model took 31.53 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] images (node: images) : f32 / [...] / [1,3,?,?] @@ -1337,7 +1496,7 @@ Instance segmentation [Step 5/11] Resizing model to match image sizes and given batch [ INFO ] Model batch size: 1 [ INFO ] Reshaping model: 'images': [1,3,640,640] - [ INFO ] Reshape model took 15.28 ms + [ INFO ] Reshape model took 16.37 ms [Step 6/11] Configuring input of the model [ INFO ] Model inputs: [ INFO ] images (node: images) : u8 / [N,C,H,W] / [1,3,640,640] @@ -1345,51 +1504,60 @@ Instance segmentation [ INFO ] output0 (node: output0) : f32 / [...] / [1,116,8400] [ INFO ] output1 (node: output1) : f32 / [...] / [1,32,160,160] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 478.74 ms + [ INFO ] Compile model took 667.41 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: + [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT [ INFO ] NETWORK_NAME: torch_jit [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 12 - [ INFO ] NUM_STREAMS: 12 - [ INFO ] AFFINITY: Affinity.CORE - [ INFO ] INFERENCE_NUM_THREADS: 36 - [ INFO ] PERF_COUNT: False - [ INFO ] INFERENCE_PRECISION_HINT: - [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT - [ INFO ] EXECUTION_MODE_HINT: ExecutionMode.PERFORMANCE - [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 - [ INFO ] ENABLE_CPU_PINNING: True - [ INFO ] SCHEDULING_CORE_TYPE: SchedulingCoreType.ANY_CORE - [ INFO ] ENABLE_HYPER_THREADING: True + [ INFO ] MODEL_PRIORITY: Priority.MEDIUM + [ INFO ] MULTI_DEVICE_PRIORITIES: CPU + [ INFO ] CPU: + [ INFO ] CPU_BIND_THREAD: YES + [ INFO ] CPU_THREADS_NUM: 0 + [ INFO ] CPU_THROUGHPUT_STREAMS: 12 + [ INFO ] DEVICE_ID: + [ INFO ] DUMP_EXEC_GRAPH_AS_DOT: + [ INFO ] DYN_BATCH_ENABLED: NO + [ INFO ] DYN_BATCH_LIMIT: 0 + [ INFO ] ENFORCE_BF16: NO + [ INFO ] EXCLUSIVE_ASYNC_REQUESTS: NO + [ INFO ] NETWORK_NAME: torch_jit + [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 12 + [ INFO ] PERFORMANCE_HINT: THROUGHPUT + [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 + [ INFO ] PERF_COUNT: NO [ INFO ] EXECUTION_DEVICES: ['CPU'] [Step 9/11] Creating infer requests and preparing input tensors [ WARNING ] No input files were given for input 'images'!. This input will be filled with random values! [ INFO ] Fill input 'images' with random values [Step 10/11] Measuring performance (Start inference asynchronously, 12 inference requests, limits: 15000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 25.31 ms + [ INFO ] First inference took 26.03 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 5088 iterations - [ INFO ] Duration: 15055.01 ms + [ INFO ] Count: 4404 iterations + [ INFO ] Duration: 15067.64 ms [ INFO ] Latency: - [ INFO ] Median: 34.93 ms - [ INFO ] Average: 35.32 ms - [ INFO ] Min: 19.79 ms - [ INFO ] Max: 105.41 ms - [ INFO ] Throughput: 337.96 FPS + [ INFO ] Median: 39.77 ms + [ INFO ] Average: 40.86 ms + [ INFO ] Min: 26.84 ms + [ INFO ] Max: 106.87 ms + [ INFO ] Throughput: 292.28 FPS -Validate quantized model accuracy -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Validate quantized model accuracy `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + As we can see, there is no significant difference between ``INT8`` and float model result in a single image test. To understand how quantization influences model prediction precision, we can compare model accuracy on a dataset. -Object detection -^^^^^^^^^^^^^^^^ +Object detection `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + .. code:: ipython3 @@ -1399,7 +1567,7 @@ Object detection .. parsed-literal:: - 0%| | 0/500 [00:00`__ +------------------------------------------------------------------------------------------------------------------------------- + .. code:: ipython3 @@ -1434,7 +1603,7 @@ Instance segmentation .. parsed-literal:: - 0%| | 0/500 [00:00`__ +############################################################################################################################### -This section contains suggestions on how to additionally improve the -performance of your application using OpenVINO. + This section contains suggestions on how to +additionally improve the performance of your application using OpenVINO. -Async inference pipeline -~~~~~~~~~~~~~~~~~~~~~~~~ +Async inference pipeline `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -The key advantage of the Async API is that when a device is busy with -inference, the application can perform other tasks in parallel (for -example, populating inputs or scheduling other requests) rather than -wait for the current inference to complete first. To understand how to -perform async inference using openvino, refer to `Async API + The key advantage of the Async +API is that when a device is busy with inference, the application can +perform other tasks in parallel (for example, populating inputs or +scheduling other requests) rather than wait for the current inference to +complete first. To understand how to perform async inference using +openvino, refer to `Async API tutorial <115-async-api-with-output.html>`__ -Integration preprocessing to model -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Integration preprocessing to model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Preprocessing API enables making preprocessing a part of the model reducing application code and dependency on additional image processing @@ -1505,8 +1676,9 @@ The integration process consists of the following steps: 3. Describe preprocessing steps. 4. Integrating Steps into a Model. -Initialize PrePostProcessing API -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +Initialize PrePostProcessing API `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + The ``openvino.preprocess.PrePostProcessor`` class enables specifying preprocessing and postprocessing steps for a model. @@ -1517,16 +1689,17 @@ preprocessing and postprocessing steps for a model. ppp = PrePostProcessor(quantized_det_model) -Define input data format -^^^^^^^^^^^^^^^^^^^^^^^^ +Define input data format `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- -To address particular input of a model/preprocessor, the -``input(input_id)`` method, where ``input_id`` is a positional index or -input tensor name for input in ``model.inputs``, if a model has a single -input, ``input_id`` can be omitted. After reading the image from the -disc, it contains U8 pixels in the ``[0, 255]`` range and is stored in -the ``NHWC`` layout. To perform a preprocessing conversion, we should -provide this to the tensor description. + To address particular input of +a model/preprocessor, the ``input(input_id)`` method, where ``input_id`` +is a positional index or input tensor name for input in +``model.inputs``, if a model has a single input, ``input_id`` can be +omitted. After reading the image from the disc, it contains U8 pixels in +the ``[0, 255]`` range and is stored in the ``NHWC`` layout. To perform +a preprocessing conversion, we should provide this to the tensor +description. .. code:: ipython3 @@ -1538,14 +1711,15 @@ provide this to the tensor description. To perform layout conversion, we also should provide information about layout expected by model -Describe preprocessing steps -^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +Describe preprocessing steps `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + Our preprocessing function contains the following steps: -* Convert the data type from ``U8`` to ``FP32``. -* Convert the data layout from ``NHWC`` to ``NCHW`` format. -* Normalize each pixel by dividing on scale factor 255. +- Convert the data type from ``U8`` to ``FP32``. +- Convert the data layout from ``NHWC`` to ``NCHW`` format. +- Normalize each pixel by dividing on scale factor 255. ``ppp.input(input_id).preprocess()`` is used for defining a sequence of preprocessing steps: @@ -1569,8 +1743,9 @@ preprocessing steps: -Integrating Steps into a Model -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +Integrating Steps into a Model `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + Once the preprocessing steps have been finished, the model can be finally built. Additionally, we can save a completed model to OpenVINO @@ -1604,7 +1779,7 @@ device. Now, we can skip these preprocessing steps in detect function: return detections - compiled_model = core.compile_model(quantized_model_with_preprocess, device) + compiled_model = core.compile_model(quantized_model_with_preprocess, device.value) input_image = np.array(Image.open(IMAGE_PATH)) detections = detect_without_preprocess(input_image, compiled_model)[0] image_with_boxes = draw_results(detections, input_image, label_map) @@ -1614,12 +1789,13 @@ device. Now, we can skip these preprocessing steps in detect function: -.. image:: 230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_85_0.png +.. image:: 230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_91_0.png -Live demo -~~~~~~~~~ +Live demo `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The following code runs model inference on a video: @@ -1631,7 +1807,7 @@ The following code runs model inference on a video: # Main processing function to run object detection. - def run_object_detection(source=0, flip=False, use_popup=False, skip_first_frames=0, model=det_model, device=device): + def run_object_detection(source=0, flip=False, use_popup=False, skip_first_frames=0, model=det_model, device="AUTO"): player = None if device != "CPU": model.reshape({0: [1, 3, 640, 640]}) @@ -1726,11 +1902,13 @@ The following code runs model inference on a video: if use_popup: cv2.destroyAllWindows() -Run -~~~ +Run `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Run Live Object Detection and Segmentation `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- -Run Live Object Detection and Segmentation -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ Use a webcam as the video input. By default, the primary webcam is set with \ ``source=0``. If you have multiple webcams, each one will be @@ -1743,7 +1921,7 @@ set \ ``use_popup=True``. notebook on a computer with a webcam. If you run the notebook on a remote server (for example, in Binder or Google Colab service), the webcam will not work. By default, the lower cell will run model - inferece on a video file. If you want to try live inference on your + inference on a video file. If you want to try live inference on your webcam set ``WEBCAM_INFERENCE = True`` Run the object detection: @@ -1757,18 +1935,45 @@ Run the object detection: else: VIDEO_SOURCE = 'https://storage.openvinotoolkit.org/repositories/openvino_notebooks/data/data/video/people.mp4' +.. code:: ipython3 + + device + + + .. parsed-literal:: - 'data/people.mp4' already exists. + Dropdown(description='Device:', index=2, options=('CPU', 'GPU', 'AUTO'), value='AUTO') + .. code:: ipython3 - run_object_detection(source=VIDEO_SOURCE, flip=True, use_popup=False, model=det_ov_model, device="AUTO") + run_object_detection(source=VIDEO_SOURCE, flip=True, use_popup=False, model=det_ov_model, device=device.value) + + + +.. image:: 230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_97_0.png + + +.. parsed-literal:: + + Source ended + Run instance segmentation: .. code:: ipython3 - run_object_detection(source=VIDEO_SOURCE, flip=True, use_popup=False, model=seg_ov_model, device="AUTO") + run_object_detection(source=VIDEO_SOURCE, flip=True, use_popup=False, model=seg_ov_model, device=device.value) + + + +.. image:: 230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_99_0.png + + +.. parsed-literal:: + + Source ended + diff --git a/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_12_1.png b/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_13_1.png similarity index 100% rename from docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_12_1.png rename to docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_13_1.png diff --git a/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_14_1.png b/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_14_1.png deleted file mode 100644 index 96939f0cf40..00000000000 --- a/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_14_1.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:91b5000d2b95f3dfb088dc1939ab379e951a5ea3f52242aa8414d33ee60c67db -size 733206 diff --git a/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_15_1.png b/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_15_1.png new file mode 100644 index 00000000000..be71ac65481 --- /dev/null +++ b/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_15_1.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:c01f6059215cb2ce6b7d9708b6d26bea27ec1f4eeff776b2223acc7715595eb2 +size 733379 diff --git a/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_24_0.png b/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_27_0.png similarity index 100% rename from docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_24_0.png rename to docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_27_0.png diff --git a/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_26_0.png b/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_29_0.png similarity index 100% rename from docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_26_0.png rename to docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_29_0.png diff --git a/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_54_0.png b/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_54_0.png deleted file mode 100644 index b8f46dc10bb..00000000000 --- a/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_54_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:16e78b19e2d4d81d71b81bc174b358532ff8b52a22590100e106e08a5396af8d -size 931323 diff --git a/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_56_0.png b/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_56_0.png deleted file mode 100644 index dd2905a32f2..00000000000 --- a/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_56_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:48f09214f8e26bcf5f66689c8173bdfc4c516e7b90ae97d40bc1cc2ec67b0e35 -size 912442 diff --git a/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_59_0.png b/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_59_0.png new file mode 100644 index 00000000000..975c9b09939 --- /dev/null +++ b/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_59_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:9fc131c6290d40ee9eaa328a07b56bd84951b2d8d4699a2e6577fe706c66062d +size 930875 diff --git a/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_61_0.png b/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_61_0.png new file mode 100644 index 00000000000..eadb7c4d7f5 --- /dev/null +++ b/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_61_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:2222119ec8c50cca18983fd246ef71351ce36524102cfa6fadb1b387304e9264 +size 912345 diff --git a/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_85_0.png b/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_85_0.png deleted file mode 100644 index b8f46dc10bb..00000000000 --- a/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_85_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:16e78b19e2d4d81d71b81bc174b358532ff8b52a22590100e106e08a5396af8d -size 931323 diff --git a/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_91_0.png b/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_91_0.png new file mode 100644 index 00000000000..975c9b09939 --- /dev/null +++ b/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_91_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:9fc131c6290d40ee9eaa328a07b56bd84951b2d8d4699a2e6577fe706c66062d +size 930875 diff --git a/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_97_0.png b/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_97_0.png new file mode 100644 index 00000000000..d86fabc033a --- /dev/null +++ b/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_97_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:b19e014c1ee03adc8fedc0932c528157c6c82f1766ba7570ee20ad704bee5784 +size 492170 diff --git a/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_99_0.png b/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_99_0.png new file mode 100644 index 00000000000..72b77ba095f --- /dev/null +++ b/docs/notebooks/230-yolov8-optimization-with-output_files/230-yolov8-optimization-with-output_99_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:87bfc5fa32aff4adefd85a6a3e47170fb94ae8d7e66168024aa3f623d583a35e +size 497983 diff --git a/docs/notebooks/230-yolov8-optimization-with-output_files/index.html b/docs/notebooks/230-yolov8-optimization-with-output_files/index.html index 9ac7b47c9a9..01c2ac260a0 100644 --- a/docs/notebooks/230-yolov8-optimization-with-output_files/index.html +++ b/docs/notebooks/230-yolov8-optimization-with-output_files/index.html @@ -1,13 +1,15 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/230-yolov8-optimization-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/230-yolov8-optimization-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/230-yolov8-optimization-with-output_files/


../
-230-yolov8-optimization-with-output_12_1.png       12-Jul-2023 00:11              909775
-230-yolov8-optimization-with-output_14_1.png       12-Jul-2023 00:11              733206
-230-yolov8-optimization-with-output_24_0.png       12-Jul-2023 00:11              931247
-230-yolov8-optimization-with-output_26_0.png       12-Jul-2023 00:11              913676
-230-yolov8-optimization-with-output_54_0.png       12-Jul-2023 00:11              931323
-230-yolov8-optimization-with-output_56_0.png       12-Jul-2023 00:11              912442
-230-yolov8-optimization-with-output_85_0.png       12-Jul-2023 00:11              931323
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/230-yolov8-optimization-with-output_files/


../
+230-yolov8-optimization-with-output_13_1.png       16-Aug-2023 01:31              909775
+230-yolov8-optimization-with-output_15_1.png       16-Aug-2023 01:31              733379
+230-yolov8-optimization-with-output_27_0.png       16-Aug-2023 01:31              931247
+230-yolov8-optimization-with-output_29_0.png       16-Aug-2023 01:31              913676
+230-yolov8-optimization-with-output_59_0.png       16-Aug-2023 01:31              930875
+230-yolov8-optimization-with-output_61_0.png       16-Aug-2023 01:31              912345
+230-yolov8-optimization-with-output_91_0.png       16-Aug-2023 01:31              930875
+230-yolov8-optimization-with-output_97_0.png       16-Aug-2023 01:31              492170
+230-yolov8-optimization-with-output_99_0.png       16-Aug-2023 01:31              497983
 

diff --git a/docs/notebooks/231-instruct-pix2pix-image-editing-with-output.rst b/docs/notebooks/231-instruct-pix2pix-image-editing-with-output.rst index d58c1aeed09..ac325b2bf11 100644 --- a/docs/notebooks/231-instruct-pix2pix-image-editing-with-output.rst +++ b/docs/notebooks/231-instruct-pix2pix-image-editing-with-output.rst @@ -1,6 +1,8 @@ Image Editing with InstructPix2Pix and OpenVINO =============================================== +.. _top: + The InstructPix2Pix is a conditional diffusion model that edits images based on written instructions provided by the user. Generative image editing models traditionally target a single editing task like style @@ -24,11 +26,26 @@ model using OpenVINO. Notebook contains the following steps: 1. Convert PyTorch models to ONNX format. -2. Convert ONNX models to OpenVINO IR format, using Model Optimizer tool. +2. Convert ONNX models to OpenVINO IR format, using model conversion + API. 3. Run InstructPix2Pix pipeline with OpenVINO. -Prerequisites -------------- + +**Table of contents**: + +- `Prerequisites <#prerequisites>`__ +- `Create Pytorch Models pipeline <#create-pytorch-models-pipeline>`__ +- `Convert Models to OpenVINO IR <#convert-models-to-openvino-ir>`__ + + - `Text Encoder <#text-encoder>`__ + - `VAE <#vae>`__ + - `Unet <#unet>`__ + +- `Prepare Inference Pipeline <#prepare-inference-pipeline>`__ + +Prerequisites `⇑ <#top>`__ +############################################################################################################################### + Install necessary packages @@ -89,8 +106,9 @@ Install necessary packages [notice] To update, run: pip install --upgrade pip -Create Pytorch Models pipeline ------------------------------- +Create Pytorch Models pipeline `⇑ <#top>`__ +############################################################################################################################### + ``StableDiffusionInstructPix2PixPipeline`` is an end-to-end inference pipeline that you can use to edit images from text instructions with @@ -126,8 +144,9 @@ First, we load the pre-trained weights of all components of the model. Fetching 15 files: 0%| | 0/15 [00:00`__ +############################################################################################################################### + OpenVINO supports PyTorch through export to the ONNX format. We will use ``torch.onnx.export`` function for obtaining an ONNX model. For more @@ -152,14 +171,17 @@ in a separate The model consists of three important parts: -* Text Encoder - to create conditions from a text prompt. -* Unet - for step-by-step denoising latent image representation. -* Autoencoder (VAE) - to encode the initial image to latent space for starting the denoising process and decoding latent space to image, when denoising is complete. +- Text Encoder - to create conditions from a text prompt. +- Unet - for step-by-step denoising latent image representation. +- Autoencoder (VAE) - to encode the initial image to latent space for + starting the denoising process and decoding latent space to image, + when denoising is complete. Let us convert each part. -Text Encoder -~~~~~~~~~~~~ +Text Encoder `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The text-encoder is responsible for transforming the input prompt, for example, “a photo of an astronaut riding a horse” into an embedding @@ -237,8 +259,9 @@ hidden states. You will use ``opset_version=14``, since model contains Text encoder will be loaded from text_encoder.xml -VAE -~~~ +VAE `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The VAE model consists of two parts: an encoder and a decoder. @@ -356,14 +379,17 @@ into two independent models. VAE decoder successfully converted to IR -Unet -~~~~ +Unet `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The Unet model has three inputs: -* ``scaled_latent_model_input`` - the latent image sample from previous step. Generation process has not been started yet, so you will use random noise. -* ``timestep`` - a current scheduler step. -* ``text_embeddings`` - a hidden state of the text encoder. +- ``scaled_latent_model_input`` - the latent image sample from previous + step. Generation process has not been started yet, so you will use + random noise. +- ``timestep`` - a current scheduler step. +- ``text_embeddings`` - a hidden state of the text encoder. Model predicts the ``sample`` state for the next step. @@ -420,8 +446,9 @@ Model predicts the ``sample`` state for the next step. Unet successfully loaded from unet.xml -Prepare Inference Pipeline --------------------------- +Prepare Inference Pipeline `⇑ <#top>`__ +############################################################################################################################### + Putting it all together, let us now take a closer look at how the model inference works by illustrating the logical flow. @@ -903,8 +930,20 @@ decoder part of the variational auto encoder. Model tokenizer and scheduler are also important parts of the pipeline. Let us define them and put all components together. Additionally, you -can provide device, for example, replace ``AUTO`` with ``GPU`` for -running model inference on GPU. +can provide device selecting one from available in dropdown list. + +.. code:: ipython3 + + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device .. code:: ipython3 @@ -913,7 +952,7 @@ running model inference on GPU. tokenizer = CLIPTokenizer.from_pretrained('openai/clip-vit-large-patch14') scheduler = EulerAncestralDiscreteScheduler.from_config(scheduler_config) - ov_pipe = OVInstructPix2PixPipeline(tokenizer, scheduler, core, TEXT_ENCODER_OV_PATH, VAE_ENCODER_OV_PATH, UNET_OV_PATH, VAE_DECODER_OV_PATH, device="AUTO") + ov_pipe = OVInstructPix2PixPipeline(tokenizer, scheduler, core, TEXT_ENCODER_OV_PATH, VAE_ENCODER_OV_PATH, UNET_OV_PATH, VAE_DECODER_OV_PATH, device=device.value) Now, you are ready to define editing instructions and an image for running the inference pipeline. You can find example results generated @@ -922,15 +961,11 @@ by the model on this need inspiration. Optionally, you can also change the random generator seed for latent state initialization and number of steps. -.. note:: - - Consider increasing ``steps`` to get more precise results. A suggested value is ``100``, but it will take more time to process. - + **Note**: Consider increasing ``steps`` to get more precise results. + A suggested value is ``100``, but it will take more time to process. .. code:: ipython3 - import ipywidgets as widgets - style = {'description_width': 'initial'} text_prompt = widgets.Text(value=" Make it in galaxy", description='your text') num_steps = widgets.IntSlider(min=1, max=100, value=10, description='steps:') @@ -997,7 +1032,7 @@ generation. -.. image:: 231-instruct-pix2pix-image-editing-with-output_files/231-instruct-pix2pix-image-editing-with-output_23_0.png +.. image:: 231-instruct-pix2pix-image-editing-with-output_files/231-instruct-pix2pix-image-editing-with-output_25_0.png Nice. As you can see, the picture has quite a high definition 🔥. diff --git a/docs/notebooks/231-instruct-pix2pix-image-editing-with-output_files/231-instruct-pix2pix-image-editing-with-output_23_0.png b/docs/notebooks/231-instruct-pix2pix-image-editing-with-output_files/231-instruct-pix2pix-image-editing-with-output_25_0.png similarity index 100% rename from docs/notebooks/231-instruct-pix2pix-image-editing-with-output_files/231-instruct-pix2pix-image-editing-with-output_23_0.png rename to docs/notebooks/231-instruct-pix2pix-image-editing-with-output_files/231-instruct-pix2pix-image-editing-with-output_25_0.png diff --git a/docs/notebooks/231-instruct-pix2pix-image-editing-with-output_files/index.html b/docs/notebooks/231-instruct-pix2pix-image-editing-with-output_files/index.html index c506b05fc39..61aac7409d2 100644 --- a/docs/notebooks/231-instruct-pix2pix-image-editing-with-output_files/index.html +++ b/docs/notebooks/231-instruct-pix2pix-image-editing-with-output_files/index.html @@ -1,7 +1,7 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/231-instruct-pix2pix-image-editing-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/231-instruct-pix2pix-image-editing-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/231-instruct-pix2pix-image-editing-with-output_files/


../
-231-instruct-pix2pix-image-editing-with-output_..> 12-Jul-2023 00:11             2122470
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/231-instruct-pix2pix-image-editing-with-output_files/


../
+231-instruct-pix2pix-image-editing-with-output_..> 16-Aug-2023 01:31             2122470
 

diff --git a/docs/notebooks/232-clip-language-saliency-map-with-output.rst b/docs/notebooks/232-clip-language-saliency-map-with-output.rst index a5dca1f4db7..728baafe9e4 100644 --- a/docs/notebooks/232-clip-language-saliency-map-with-output.rst +++ b/docs/notebooks/232-clip-language-saliency-map-with-output.rst @@ -88,9 +88,16 @@ Initial Implementation with Transformers and Pytorch .. code:: ipython3 # Install requirements - !pip install -q 'openvino-dev>=2023.0.0' + !pip install -q "openvino-dev>=2023.0.0" !pip install -q onnx transformers torch + +.. parsed-literal:: + + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + + .. code:: ipython3 from pathlib import Path @@ -107,7 +114,7 @@ Initial Implementation with Transformers and Pytorch To get the CLIP model, you will use the ``transformers`` library and the official ``openai/clip-vit-base-patch16`` from OpenAI. You can use any -CLIP model from the Huggingface Hub by simply replacing a model +CLIP model from the HuggingFace Hub by simply replacing a model checkpoint in the cell below. There are several preprocessing steps required to get text and image @@ -126,10 +133,10 @@ steps. .. parsed-literal:: - 2023-07-11 23:27:42.102607: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. - 2023-07-11 23:27:42.133276: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + 2023-07-18 23:28:44.655634: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-07-18 23:28:44.687925: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. - 2023-07-11 23:27:42.701655: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + 2023-07-18 23:28:45.260957: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT Let us write helper functions first. You will generate crop coordinates @@ -174,12 +181,13 @@ formula above. ) -> Union[np.ndarray, torch.Tensor]: return one @ other.T / (np.linalg.norm(one) * np.linalg.norm(other)) -Parameters to be defined: - -- ``n_iters`` - number of times the procedure will be repeated. Larger is better, but will require more time to inference -- ``min_crop_size`` - minimum size of the crop window. A smaller size will increase the resolution of the saliency map but may require more iterations -- ``query`` - text that will be used to query the image -- ``image`` - the actual image that will be queried. You will download the image from a link +Parameters to be defined: - ``n_iters`` - number of times the procedure +will be repeated. Larger is better, but will require more time to +inference - ``min_crop_size`` - minimum size of the crop window. A +smaller size will increase the resolution of the saliency map but may +require more iterations - ``query`` - text that will be used to query +the image - ``image`` - the actual image that will be queried. You will +download the image from a link The image at the beginning was acquired with ``n_iters=2000`` and ``min_crop_size=50``. You will start with the lower number of inferences @@ -383,17 +391,15 @@ information on ONNX conversion. .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/clip/modeling_clip.py:284: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-453/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/clip/modeling_clip.py:286: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if attn_weights.size() != (bsz * self.num_heads, tgt_len, src_len): - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/clip/modeling_clip.py:324: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-453/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/clip/modeling_clip.py:326: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if attn_output.size() != (bsz * self.num_heads, tgt_len, self.head_dim): - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/clip/modeling_clip.py:684: TracerWarning: torch.tensor results are registered as constants in the trace. You can safely ignore this warning if you use this function to create tensors out of constant variables that would be the same every time you call this function. In any other case, this might cause the trace to be incorrect. - mask = torch.full((tgt_len, tgt_len), torch.tensor(torch.finfo(dtype).min, device=device), device=device) - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/clip/modeling_clip.py:292: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-453/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/clip/modeling_clip.py:294: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if causal_attention_mask.size() != (bsz, 1, tgt_len, src_len): - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/clip/modeling_clip.py:301: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-453/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/clip/modeling_clip.py:303: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if attention_mask.size() != (bsz, 1, tgt_len, src_len): - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/symbolic_opset9.py:5408: UserWarning: Exporting aten::index operator of advanced indexing in opset 14 is achieved by combination of multiple ONNX operators, including Reshape, Transpose, Concat, and Gather. If indices include negative values, the exported graph will produce incorrect results. + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-453/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/symbolic_opset9.py:5408: UserWarning: Exporting aten::index operator of advanced indexing in opset 14 is achieved by combination of multiple ONNX operators, including Reshape, Transpose, Concat, and Gather. If indices include negative values, the exported graph will produce incorrect results. warnings.warn( @@ -476,14 +482,42 @@ Inference with OpenVINO™ from openvino.runtime import Core - core = Core() text_model = core.read_model(text_model_path) image_model = core.read_model(image_model_path) + +Select inference device +~~~~~~~~~~~~~~~~~~~~~~~ + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets - text_model = core.compile_model(model=text_model, device_name="CPU") - image_model = core.compile_model(model=image_model, device_name="CPU") + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + text_model = core.compile_model(model=text_model, device_name=device.value) + image_model = core.compile_model(model=image_model, device_name=device.value) OpenVINO supports ``numpy.ndarray`` as an input type, so you change the ``return_tensors`` to ``np``. You also convert a transformers’ @@ -528,11 +562,11 @@ the inference process is mostly similar. -.. image:: 232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_28_1.png +.. image:: 232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_31_1.png -Accelerate Inference with AsyncInferQueue ------------------------------------------ +Accelerate Inference with ``AsyncInferQueue`` +--------------------------------------------- Up until now, the pipeline was synchronous, which means that the data preparation, model input population, model inference, and output @@ -570,7 +604,7 @@ performance hint. image_model = core.compile_model( model=image_model, - device_name="CPU", + device_name=device.value, config={"PERFORMANCE_HINT":"THROUGHPUT"}, ) @@ -588,11 +622,9 @@ performance hint. saliency_map = np.zeros((y_dim, x_dim)) Your callback should do the same thing that you did after inference in -the sync mode: - -- Pull the image embeddings from an inference request. -- Compute cosine similarity between text and image embeddings. -- Update saliency map based. +the sync mode: - Pull the image embeddings from an inference request. - +Compute cosine similarity between text and image embeddings. - Update +saliency map based. If you do not change the progress bar, it will show the progress of pushing data to the inference queue. To track the actual progress, you @@ -657,7 +689,7 @@ should pass a progress bar object and call ``update`` method after -.. image:: 232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_34_1.png +.. image:: 232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_37_1.png Pack the Pipeline into a Function @@ -792,12 +824,15 @@ What To Do Next --------------- Now that you have a convenient interface and accelerated inference, you -can explore the CLIP capabilities further. For example: - -- Can CLIP read? Can it detect text regions in general and specific words on the -image? -- Which famous people and places does CLIP know? - Can CLIP identify places on -a map? Or planets, stars, and constellations? -- Explore different CLIP models from Huggingface Hub: just change the ``model_checkpoint`` at the beginning of the notebook. -- Add batch processing to the pipeline: modify ``get_random_crop_params``, ``get_cropped_image`` and ``update_saliency_map`` functions to process multiple crop images at once and accelerate the pipeline even more. -- Optimize models with `NNCF `__ to get further acceleration. +can explore the CLIP capabilities further. For example: - Can CLIP read? +Can it detect text regions in general and specific words on the image? - +Which famous people and places does CLIP know? - Can CLIP identify +places on a map? Or planets, stars, and constellations? - Explore +different CLIP models from HuggingFace Hub: just change the +``model_checkpoint`` at the beginning of the notebook. - Add batch +processing to the pipeline: modify ``get_random_crop_params``, +``get_cropped_image`` and ``update_saliency_map`` functions to process +multiple crop images at once and accelerate the pipeline even more. - +Optimize models with +`NNCF `__ +to get further acceleration. diff --git a/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_15_0.png b/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_15_0.png index 4586808a227..d9334001b5c 100644 --- a/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_15_0.png +++ b/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_15_0.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:0f9111cec8f1c4110326237e32ec210f522956f3de929cb48e43c22bd2777958 -size 78374 +oid sha256:de55a51d782774cc789cdf9d8759541f9e5aabef78730786641c2affc8bfb09e +size 73946 diff --git a/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_17_0.png b/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_17_0.png index a15bd4b19f6..69ac5c8055b 100644 --- a/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_17_0.png +++ b/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_17_0.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:871b03a85e36eab569154bcb8997715bb55a9fca5549d4de4f59c79af98fa3c9 -size 501300 +oid sha256:369238063ffcca3d69de022cae0f47ef5dab6c9440f15a60cbcf325ba4543e9e +size 499941 diff --git a/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_19_1.png b/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_19_1.png index dc4827afd3b..286025078a3 100644 --- a/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_19_1.png +++ b/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_19_1.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:2caf34efdfbdccaa62dcb53fb70017f80458353c5e600774fcb367e9fb8bf890 -size 503710 +oid sha256:38cfca02bb2ce94e34e3b1e8302a4deb274696b21b3c144c502f9019441530e5 +size 502742 diff --git a/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_28_1.png b/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_28_1.png deleted file mode 100644 index c3a2550bc2e..00000000000 --- a/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_28_1.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:23ee60550875094a9185f7adb9945b8418791dc8a4fb9feaeda34d13b3c94857 -size 496081 diff --git a/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_31_1.png b/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_31_1.png new file mode 100644 index 00000000000..53305025624 --- /dev/null +++ b/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_31_1.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:dfd7bb983b20aa47ced3e0c311860da50090522d81d2e0f75fb92cd09e1648cb +size 501301 diff --git a/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_34_1.png b/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_34_1.png deleted file mode 100644 index 8e44afc0c76..00000000000 --- a/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_34_1.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:698efb32124f868b8927f56d3acf62ce37a137f295bd5ef1b94b4fb815b972b4 -size 500537 diff --git a/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_37_1.png b/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_37_1.png new file mode 100644 index 00000000000..9544a8a17ec --- /dev/null +++ b/docs/notebooks/232-clip-language-saliency-map-with-output_files/232-clip-language-saliency-map-with-output_37_1.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:2683fdc809e1aaa072f65543a725c245870345c895b17cd430bac2c55e3f2409 +size 496940 diff --git a/docs/notebooks/232-clip-language-saliency-map-with-output_files/index.html b/docs/notebooks/232-clip-language-saliency-map-with-output_files/index.html deleted file mode 100644 index 69a334bcb0a..00000000000 --- a/docs/notebooks/232-clip-language-saliency-map-with-output_files/index.html +++ /dev/null @@ -1,11 +0,0 @@ - -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/232-clip-language-saliency-map-with-output_files/ - -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/232-clip-language-saliency-map-with-output_files/


../
-232-clip-language-saliency-map-with-output_15_0..> 12-Jul-2023 00:11               78374
-232-clip-language-saliency-map-with-output_17_0..> 12-Jul-2023 00:11              501300
-232-clip-language-saliency-map-with-output_19_1..> 12-Jul-2023 00:11              503710
-232-clip-language-saliency-map-with-output_28_1..> 12-Jul-2023 00:11              496081
-232-clip-language-saliency-map-with-output_34_1..> 12-Jul-2023 00:11              500537
-

- diff --git a/docs/notebooks/233-blip-visual-language-processing-with-output.rst b/docs/notebooks/233-blip-visual-language-processing-with-output.rst index 7f3343fee05..2637f314bf1 100644 --- a/docs/notebooks/233-blip-visual-language-processing-with-output.rst +++ b/docs/notebooks/233-blip-visual-language-processing-with-output.rst @@ -1,6 +1,8 @@ Visual Question Answering and Image Captioning using BLIP and OpenVINO ====================================================================== +.. _top: + Humans perceive the world through vision and language. A longtime goal of AI is to build intelligent agents that can understand the world through vision and language inputs to communicate with humans through @@ -22,21 +24,43 @@ The tutorial consists of the following parts: 2. Convert the BLIP model to OpenVINO IR. 3. Run visual question answering and image captioning with OpenVINO. -Background ----------- +**Table of contents**: -Visual language processing is a branch of an artificial intelligence -that focuses on creating algorithms designed to enable computers to more +- `Background <#background>`__ + + - `Image Captioning <#image-captioning>`__ + - `Visual Question Answering <#visual-question-answering>`__ + +- `Instantiate Model <#instantiate-model>`__ +- `Convert Models to OpenVINO IR <#convert-models-to-openvino-ir>`__ + + - `Vision Model <#vision-model>`__ + - `Text Encoder <#text-encoder>`__ + - `Text Decoder <#text-decoder>`__ + +- `Run OpenVINO Model <#run-openvino-model>`__ + + - `Prepare Inference Pipeline <#prepare-inference-pipeline>`__ + - `Select inference device <#select-inference-device>`__ + - `Image Captioning <#image-captioning>`__ + - `Question Answering <#question-answering>`__ + +Background `⇑ <#top>`__ +############################################################################################################################### + + +Visual language processing is a branch of artificial intelligence that +focuses on creating algorithms designed to enable computers to more accurately understand images and their content. Popular tasks include: -* **Text to Image Retrieval** - a semantic task -that aims to find the most relevant image for a given text description. -* **Image Captioning** - a semantic task that aims to provide a text -description for image content. -* **Visual Question Answering** - a -semantic task that aims to answer questions based on image content. +- **Text to Image Retrieval** - a semantic task that aims to find the + most relevant image for a given text description. +- **Image Captioning** - a semantic task that aims to provide a text + description for image content. +- **Visual Question Answering** - a semantic task that aims to answer + questions based on image content. As shown in the diagram below, these three tasks differ in the input provided to the AI system. For text-to-image retrieval, you have a @@ -47,13 +71,14 @@ case of visual question answering, where you have a predefined question visual question answering, both the text-based question and image context are variables requested by a user. -.. image:: https://camo.githubusercontent.com/3a1ebcd9609c551f47ad4e030556dc5b15f1f5a00e09b81facbfc85152a541a6/68747470733a2f2f757365722d696d616765732e67697468756275736572636f6e74656e742e636f6d2f32393435343439392f3232313735353731372d61356235316237652d353233632d343631662d623330632d3465646266616639613133342e706e67 +|image0| This notebook does not focus on Text to Image retrieval. Instead, it considers Image Captioning and Visual Question Answering. -Image Captioning -~~~~~~~~~~~~~~~~ +Image Captioning `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Image Captioning is the task of describing the content of an image in words. This task lies at the intersection of computer vision and natural @@ -62,22 +87,21 @@ encoder-decoder framework, where an input image is encoded into an intermediate representation of the information in the image, and then decoded into a descriptive text sequence. -.. image:: https://camo.githubusercontent.com/13bad05eeae51c0318331a4bb0e30e9d1e196fa3127fafd39ae30ddee4751efe/68747470733a2f2f757365722d696d616765732e67697468756275736572636f6e74656e742e636f6d2f32393435343439392f3232313634303834372d31383638313137632d616163302d343830362d393961342d3334663231386539386262382e706e67 +|image1| -Visual Question Answering -~~~~~~~~~~~~~~~~~~~~~~~~~ +Visual Question Answering `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Visual Question Answering (VQA) is the task of answering text-based -questions about image content. +Visual Question Answering (VQA) is the task of answering text-based questions about image content. -.. image:: https://camo.githubusercontent.com/43412d73adb987a8ebf397cd10f0b0e98e2b436bb13195bcc03930c44dcb1740/68747470733a2f2f757365722d696d616765732e67697468756275736572636f6e74656e742e636f6d2f32393435343439392f3232313634313938342d33633664386232662d646430642d343330322d613464382d3066383536346663613737322e706e67 +|image2| For a better understanding of how VQA works, let us consider a traditional NLP task like Question Answering, which aims to retrieve the answer to a question from a given text input. Typically, a question answering pipeline consists of three steps: -.. image:: https://camo.githubusercontent.com/434ec750fdd38953ccf2d0dfb9c545edd9cb2681ad4ffcd00f6a66c62df9f1ca/68747470733a2f2f757365722d696d616765732e67697468756275736572636f6e74656e742e636f6d2f32393435343439392f3232313736303838312d33373866316561382d656164632d343631302d616666302d3639656361626636326666662e706e67 +|image3| 1. Question analysis - analysis of provided question in natural language form to understand the object in the question and additional context. @@ -92,70 +116,79 @@ answering pipeline consists of three steps: knowledge base, typically provided text documents or databases serve as a source of knowledge. -.. image:: https://camo.githubusercontent.com/f06359136073bc9162207c04d993755f49a55957dabeb11ebd9b9ee08638234b/68747470733a2f2f757365722d696d616765732e67697468756275736572636f6e74656e742e636f6d2f32393435343439392f3232323039343836312d33636166646639662d643730302d343734312d623663352d6662303963316134646139612e706e67 +|image4| The difference between text-based question answering and visual question answering is that an image is used as context and the knowledge base. -.. image:: https://camo.githubusercontent.com/c7f49075b23edf50a7f2ae9263ad39828226961a205b70a1c18c12576047ba85/68747470733a2f2f757365722d696d616765732e67697468756275736572636f6e74656e742e636f6d2f32393435343439392f3232323039353131382d33643538323665342d323636322d346431632d616266322d6135313566323364366436612e706e67 +|image5| Answering arbitrary questions about images is a complex problem because it requires involving a lot of computer vision sub-tasks. In the table below, you can find an example of questions and the required computer vision skills to find answers. -+--------------------+-------------------------------------------------+ -| Computer vision | Question examples | -| task | | -+====================+=================================================+ -| Object recognition | What is shown in the picture? What is it? | -+--------------------+-------------------------------------------------+ -| Object detection | Is there any object (dog, man, book) in the | -| | image? Where is … located? | -+--------------------+-------------------------------------------------+ -| Object and image | What color is an umbrella? Does this man wear | -| attribute | glasses? Is there color in the image? | -| recognition | | -+--------------------+-------------------------------------------------+ -| Scene recognition | Is it rainy? What celebration is pictured? | -+--------------------+-------------------------------------------------+ -| Object counting | How many players are there on the football | -| | field? How many steps are there on the stairs? | -+--------------------+-------------------------------------------------+ -| Activity | Is the baby crying? What is the woman cooking? | -| recognition | What are they doing? | -+--------------------+-------------------------------------------------+ -| Spatial | What is located between the sofa and the | -| relationships | armchair? What is in the bottom left corner? | -| among objects | | -+--------------------+-------------------------------------------------+ -| Commonsense | Does she have 100% vision? Does this person | -| reasoning | have children? | -+--------------------+-------------------------------------------------+ -| Knowledge-based | Is it a vegetarian pizza? | -| reasoning | | -+--------------------+-------------------------------------------------+ -| Text recognition | What is the title of the book? What is shown on | -| | the screen? | -+--------------------+-------------------------------------------------+ ++-----------------------------+----------------------------------------+ +| Computer vision task | Question examples | ++=============================+========================================+ +| Object recognition | What is shown in the picture? What is | +| | it? | ++-----------------------------+----------------------------------------+ +| Object detection | Is there any object (dog, man, book) | +| | in the image? Where is … located? | ++-----------------------------+----------------------------------------+ +| Object and image attribute | What color is an umbrella? Does this | +| recognition | man wear glasses? Is there color in | +| | the image? | ++-----------------------------+----------------------------------------+ +| Scene recognition | Is it rainy? What celebration is | +| | pictured? | ++-----------------------------+----------------------------------------+ +| Object counting | How many players are there on the | +| | football field? How many steps are | +| | there on the stairs? | ++-----------------------------+----------------------------------------+ +| Activity recognition | Is the baby crying? What is the woman | +| | cooking? What are they doing? | ++-----------------------------+----------------------------------------+ +| Spatial relationships among | What is located between the sofa and | +| objects | the armchair? What is in the bottom | +| | left corner? | ++-----------------------------+----------------------------------------+ +| Commonsense reasoning | Does she have 100% vision? Does this | +| | person have children? | ++-----------------------------+----------------------------------------+ +| Knowledge-based reasoning | Is it a vegetarian pizza? | ++-----------------------------+----------------------------------------+ +| Text recognition | What is the title of the book? What is | +| | shown on the screen? | ++-----------------------------+----------------------------------------+ -There are a lot of applications for visual question answering: +There are a lot of applications for visual question answering: -* Aid Visually Impaired Persons: VQA models can be used to reduce barriers for -visually impaired people by helping them get information about images -from the web and the real world. -* Education: VQA models can be used to -improve visitor experiences at museums by enabling observers to directly -ask questions they are interested in or to bring more interactivity to -schoolbooks for children interested in acquiring specific knowledge. -* E-commerce: VQA models can retrieve information about products using -photos from online stores. -* Independent expert assessment: VQA models -can be provide objective assessments in sports competitions, medical -diagnosis, and forensic examination. +- Aid Visually Impaired Persons: VQA models can be used to reduce + barriers for visually impaired people by helping them get information + about images from the web and the real world. +- Education: VQA models can be used to improve visitor experiences at + museums by enabling observers to directly ask questions they are + interested in or to bring more interactivity to schoolbooks for + children interested in acquiring specific knowledge. +- E-commerce: VQA models can retrieve information about products using + photos from online stores. +- Independent expert assessment: VQA models can be provide objective + assessments in sports competitions, medical diagnosis, and forensic + examination. + +.. |image0| image:: https://user-images.githubusercontent.com/29454499/221755717-a5b51b7e-523c-461f-b30c-4edbfaf9a134.png +.. |image1| image:: https://user-images.githubusercontent.com/29454499/221640847-1868117c-aac0-4806-99a4-34f218e98bb8.png +.. |image2| image:: https://user-images.githubusercontent.com/29454499/221641984-3c6d8b2f-dd0d-4302-a4d8-0f8564fca772.png +.. |image3| image:: https://user-images.githubusercontent.com/29454499/221760881-378f1ea8-eadc-4610-aff0-69ecabf62fff.png +.. |image4| image:: https://user-images.githubusercontent.com/29454499/222094861-3cafdf9f-d700-4741-b6c5-fb09c1a4da9a.png +.. |image5| image:: https://user-images.githubusercontent.com/29454499/222095118-3d5826e4-2662-4d1c-abf2-a515f23d6d6a.png + +Instantiate Model `⇑ <#top>`__ +############################################################################################################################### -Instantiate Model ------------------ The BLIP model was proposed in the `BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and @@ -171,26 +204,26 @@ generation capabilities, BLIP introduces a multimodal mixture of an encoder-decoder and a multi-task model which can operate in one of the three modes: -* **Unimodal encoders**, which separately encode images -and text. The image encoder is a vision transformer. The text encoder is -the same as BERT. -* **Image-grounded text encoder**, which injects -visual information by inserting a cross-attention layer between the -self-attention layer and the feed-forward network for each transformer -block of the text encoder. -* **Image-grounded text decoder**, which -replaces the bi-directional self-attention layers in the text encoder -with causal self-attention layers. +- **Unimodal encoders**, which separately encode images and text. The + image encoder is a vision transformer. The text encoder is the same + as BERT. +- **Image-grounded text encoder**, which injects visual information by + inserting a cross-attention layer between the self-attention layer + and the feed-forward network for each transformer block of the text + encoder. +- **Image-grounded text decoder**, which replaces the bi-directional + self-attention layers in the text encoder with causal self-attention + layers. More details about the model can be found in the `research -paper `__, `Salesforces +paper `__, `Salesforce blog `__, `GitHub repo `__ and `Hugging Face model documentation `__. In this tutorial, you will use the -```blip-vqa-base`` `__ +`blip-vqa-base `__ model available for download from `Hugging Face `__. The same actions are also applicable to other similar models from the BLIP family. Although this model class @@ -209,24 +242,25 @@ text and vision modalities and postprocessing of generation results. .. parsed-literal:: - Requirement already satisfied: transformers>=4.26.0 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (4.30.2) - Requirement already satisfied: filelock in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (3.12.2) - Requirement already satisfied: huggingface-hub<1.0,>=0.14.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (0.16.4) - Requirement already satisfied: numpy>=1.17 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (1.23.5) - Requirement already satisfied: packaging>=20.0 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (23.1) - Requirement already satisfied: pyyaml>=5.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (6.0) - Requirement already satisfied: regex!=2019.12.17 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (2023.6.3) - Requirement already satisfied: requests in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (2.31.0) - Requirement already satisfied: tokenizers!=0.11.3,<0.14,>=0.11.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (0.13.3) - Requirement already satisfied: safetensors>=0.3.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (0.3.1) - Requirement already satisfied: tqdm>=4.27 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (4.65.0) - Requirement already satisfied: fsspec in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from huggingface-hub<1.0,>=0.14.1->transformers>=4.26.0) (2023.6.0) - Requirement already satisfied: typing-extensions>=3.7.4.3 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from huggingface-hub<1.0,>=0.14.1->transformers>=4.26.0) (4.7.1) - Requirement already satisfied: charset-normalizer<4,>=2 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests->transformers>=4.26.0) (3.2.0) - Requirement already satisfied: idna<4,>=2.5 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests->transformers>=4.26.0) (3.4) - Requirement already satisfied: urllib3<3,>=1.21.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests->transformers>=4.26.0) (1.26.16) - Requirement already satisfied: certifi>=2017.4.17 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests->transformers>=4.26.0) (2023.5.7) - + Requirement already satisfied: transformers>=4.26.0 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (4.31.0) + Requirement already satisfied: filelock in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (3.12.2) + Requirement already satisfied: huggingface-hub<1.0,>=0.14.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (0.16.4) + Requirement already satisfied: numpy>=1.17 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (1.23.5) + Requirement already satisfied: packaging>=20.0 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (23.1) + Requirement already satisfied: pyyaml>=5.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (6.0.1) + Requirement already satisfied: regex!=2019.12.17 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (2023.8.8) + Requirement already satisfied: requests in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (2.31.0) + Requirement already satisfied: tokenizers!=0.11.3,<0.14,>=0.11.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (0.13.3) + Requirement already satisfied: safetensors>=0.3.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (0.3.2) + Requirement already satisfied: tqdm>=4.27 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from transformers>=4.26.0) (4.66.1) + Requirement already satisfied: fsspec in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from huggingface-hub<1.0,>=0.14.1->transformers>=4.26.0) (2023.6.0) + Requirement already satisfied: typing-extensions>=3.7.4.3 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from huggingface-hub<1.0,>=0.14.1->transformers>=4.26.0) (4.7.1) + Requirement already satisfied: charset-normalizer<4,>=2 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests->transformers>=4.26.0) (3.2.0) + Requirement already satisfied: idna<4,>=2.5 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests->transformers>=4.26.0) (3.4) + Requirement already satisfied: urllib3<3,>=1.21.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests->transformers>=4.26.0) (1.26.16) + Requirement already satisfied: certifi>=2017.4.17 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from requests->transformers>=4.26.0) (2023.7.22) + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + .. code:: ipython3 @@ -261,10 +295,10 @@ text and vision modalities and postprocessing of generation results. .. parsed-literal:: - 2023-07-11 23:28:58.566617: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. - 2023-07-11 23:28:58.600974: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + 2023-08-15 23:34:17.871379: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-08-15 23:34:17.904962: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. - 2023-07-11 23:28:59.072595: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + 2023-08-15 23:34:18.440790: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT @@ -275,7 +309,7 @@ text and vision modalities and postprocessing of generation results. .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/generation/utils.py:1353: UserWarning: Using `max_length`'s default (20) to control the generation length. This behaviour is deprecated and will be removed from the config in v5 of Transformers -- we recommend using `max_new_tokens` to control the maximum length of the generation. + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/generation/utils.py:1369: UserWarning: Using `max_length`'s default (20) to control the generation length. This behaviour is deprecated and will be removed from the config in v5 of Transformers -- we recommend using `max_new_tokens` to control the maximum length of the generation. warnings.warn( @@ -286,7 +320,7 @@ text and vision modalities and postprocessing of generation results. .. parsed-literal:: - Processing time: 0.2072 s + Processing time: 0.2136 s .. code:: ipython3 @@ -327,11 +361,12 @@ text and vision modalities and postprocessing of generation results. -.. image:: 233-blip-visual-language-processing-with-output_files/233-blip-visual-language-processing-with-output_7_0.png +.. image:: 233-blip-visual-language-processing-with-output_files/233-blip-visual-language-processing-with-output_8_0.png -Convert Models to OpenVINO IR ------------------------------ +Convert Models to OpenVINO IR `⇑ <#top>`__ +############################################################################################################################### + OpenVINO supports PyTorch through export to the ONNX format. You will use the ``torch.onnx.export`` function for obtaining ONNX model. For @@ -345,22 +380,22 @@ example, input and output names or dynamic shapes). While ONNX models are directly supported by OpenVINO™ runtime, it can be useful to convert them to OpenVINO Intermediate Representation (IR) format to take the advantage of advanced OpenVINO optimization tools and -features. You will use OpenVINO Model Optimizer to convert the model to -IR format and compress weights to ``FP16`` format. +features. You will use model conversion API to convert the model to IR +format and compress weights to ``FP16`` format. -The model consists of three parts: +The model consists of three parts: -* vision_model - an encoder for -image representation. -* text_encoder - an encoder for input query, used -for question answering and text-to-image retrieval only. -* text_decoder - a decoder for output answer. +- vision_model - an encoder for image representation. +- text_encoder - an encoder for input query, used for question + answering and text-to-image retrieval only. +- text_decoder - a decoder for output answer. To be able to perform multiple tasks, using the same model components, you should convert each part independently. -Vision Model -~~~~~~~~~~~~ +Vision Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The vision model accepts float input tensors with the [1,3,384,384] shape, containing RGB image pixel values normalized in the [0,1] range. @@ -388,7 +423,7 @@ shape, containing RGB image pixel values normalized in the [0,1] range. if not VISION_MODEL_ONNX.exists(): with torch.no_grad(): torch.onnx.export(vision_model, inputs["pixel_values"], VISION_MODEL_ONNX, input_names=["pixel_values"]) - # convert ONNX model to IR using Model Optimizer Python API, use compress_to_fp16=True for compressing model weights to FP16 precision + # convert ONNX model to IR using model conversion Python API, use compress_to_fp16=True for compressing model weights to FP16 precision ov_vision_model = mo.convert_model(VISION_MODEL_ONNX, compress_to_fp16=True) # save model on disk for next usages serialize(ov_vision_model, str(VISION_MODEL_OV)) @@ -406,8 +441,9 @@ shape, containing RGB image pixel values normalized in the [0,1] range. Vision model successfuly converted and saved to blip_vision_model.xml -Text Encoder -~~~~~~~~~~~~ +Text Encoder `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The text encoder is used by visual question answering tasks to build a question embedding representation. It takes ``input_ids`` with a @@ -443,7 +479,7 @@ tutorial `__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The text decoder is responsible for generating the sequence of tokens to represent model output (answer to question or caption), using an image @@ -549,7 +586,7 @@ shapes. if not TEXT_DECODER_ONNX.exists(): with torch.no_grad(): torch.onnx.export(text_decoder, input_dict, TEXT_DECODER_ONNX, input_names=list(input_dict), output_names=output_names + past_key_values_outs, dynamic_axes=dynamic_axes) - # convert ONNX model to IR using Model Optimizer Python API, use compress_to_fp16=True for compressing model weights to FP16 precision + # convert ONNX model to IR using model conversion Python API, use compress_to_fp16=True for compressing model weights to FP16 precision ov_text_decoder = mo.convert_model(TEXT_DECODER_ONNX, compress_to_fp16=True) # save model on disk for next usages serialize(ov_text_decoder, str(TEXT_DECODER_OV)) @@ -560,9 +597,9 @@ shapes. .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/blip/modeling_blip_text.py:639: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/blip/modeling_blip_text.py:640: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if causal_mask.shape[1] < attention_mask.shape[1]: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/blip/modeling_blip_text.py:890: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/transformers/models/blip/modeling_blip_text.py:888: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if return_logits: @@ -600,7 +637,7 @@ layers. if not TEXT_DECODER_WITH_PAST_ONNX.exists(): with torch.no_grad(): torch.onnx.export(text_decoder, input_dict_with_past, TEXT_DECODER_WITH_PAST_ONNX, input_names=input_names_with_past, output_names=output_names + past_key_values_outs, dynamic_axes=dynamic_axes_with_past) - # convert ONNX model to IR using Model Optimizer Python API, use compress_to_fp16=True for compressing model weights to FP16 precision + # convert ONNX model to IR using model conversion Python API, use compress_to_fp16=True for compressing model weights to FP16 precision ov_text_decoder = mo.convert_model(TEXT_DECODER_WITH_PAST_ONNX, compress_to_fp16=True) # save model on disk for next usages serialize(ov_text_decoder, str(TEXT_DECODER_WITH_PAST_OV)) @@ -614,44 +651,79 @@ layers. Text decoder with past successfuly converted and saved to blip_text_decoder_with_past.xml -Run OpenVINO Model ------------------- +Run OpenVINO Model `⇑ <#top>`__ +############################################################################################################################### + + +Prepare Inference Pipeline `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Prepare Inference Pipeline -~~~~~~~~~~~~~~~~~~~~~~~~~~ As discussed before, the model consists of several blocks which can be reused for building pipelines for different tasks. In the diagram below, you can see how image captioning works: -.. image:: https://camo.githubusercontent.com/70694ac4d489ffcece4c48133e5e462c52858a9ebed6db6a6d61af2e91ae987a/68747470733a2f2f757365722d696d616765732e67697468756275736572636f6e74656e742e636f6d2f32393435343439392f3232313836353833362d61353664613036652d313936642d343439632d613564632d3431333664613661623564352e706e67 +|image01| -The visual model accepts the image preprocessed by BlipProcessor as +The visual model accepts the image preprocessed by ``BlipProcessor`` as input and produces image embeddings, which are directly passed to the text decoder for generation caption tokens. When generation is finished, -output sequence of tokens is provided to BlipProcessor for decoding to -text using a tokenizer. +output sequence of tokens is provided to ``BlipProcessor`` for decoding +to text using a tokenizer. The pipeline for question answering looks similar, but with additional question processing. In this case, image embeddings and question -tokenized by BlipProcessor are provided to the text encoder and then +tokenized by ``BlipProcessor`` are provided to the text encoder and then multimodal question embedding is passed to the text decoder for performing generation of answers. -.. image:: https://camo.githubusercontent.com/a90752e235ef0ab74382290e6adad951b16101578a98e220f3872245b8e223a7/68747470733a2f2f757365722d696d616765732e67697468756275736572636f6e74656e742e636f6d2f32393435343439392f3232313836383136372d64303038316164642d643966332d343539312d383065372d3437353363383863316430612e706e67 +|image02| The next step is implementing both pipelines using OpenVINO models. +.. |image01| image:: https://user-images.githubusercontent.com/29454499/221865836-a56da06e-196d-449c-a5dc-4136da6ab5d5.png +.. |image02| image:: https://user-images.githubusercontent.com/29454499/221868167-d0081add-d9f3-4591-80e7-4753c88c1d0a.png + .. code:: ipython3 # create OpenVINO Core object instance core = Core() + +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + # load models on device - ov_vision_model = core.compile_model(VISION_MODEL_OV) - ov_text_encoder = core.compile_model(TEXT_ENCODER_OV) - ov_text_decoder = core.compile_model(TEXT_DECODER_OV) - ov_text_decoder_with_past = core.compile_model(TEXT_DECODER_WITH_PAST_OV) + ov_vision_model = core.compile_model(VISION_MODEL_OV, device.value) + ov_text_encoder = core.compile_model(TEXT_ENCODER_OV, device.value) + ov_text_decoder = core.compile_model(TEXT_DECODER_OV, device.value) + ov_text_decoder_with_past = core.compile_model(TEXT_DECODER_WITH_PAST_OV, device.value) .. code:: ipython3 @@ -826,8 +898,9 @@ initial token for decoder work. Now, the model is ready for generation. -Image Captioning -~~~~~~~~~~~~~~~~ +Image Captioning `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -837,11 +910,12 @@ Image Captioning -.. image:: 233-blip-visual-language-processing-with-output_files/233-blip-visual-language-processing-with-output_24_0.png +.. image:: 233-blip-visual-language-processing-with-output_files/233-blip-visual-language-processing-with-output_28_0.png -Question Answering -~~~~~~~~~~~~~~~~~~ +Question Answering `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -853,7 +927,7 @@ Question Answering -.. image:: 233-blip-visual-language-processing-with-output_files/233-blip-visual-language-processing-with-output_26_0.png +.. image:: 233-blip-visual-language-processing-with-output_files/233-blip-visual-language-processing-with-output_30_0.png .. code:: ipython3 @@ -863,5 +937,5 @@ Question Answering .. parsed-literal:: - Processing time: 0.1530 + Processing time: 0.1504 diff --git a/docs/notebooks/233-blip-visual-language-processing-with-output_files/233-blip-visual-language-processing-with-output_24_0.png b/docs/notebooks/233-blip-visual-language-processing-with-output_files/233-blip-visual-language-processing-with-output_28_0.png similarity index 100% rename from docs/notebooks/233-blip-visual-language-processing-with-output_files/233-blip-visual-language-processing-with-output_24_0.png rename to docs/notebooks/233-blip-visual-language-processing-with-output_files/233-blip-visual-language-processing-with-output_28_0.png diff --git a/docs/notebooks/233-blip-visual-language-processing-with-output_files/233-blip-visual-language-processing-with-output_26_0.png b/docs/notebooks/233-blip-visual-language-processing-with-output_files/233-blip-visual-language-processing-with-output_30_0.png similarity index 100% rename from docs/notebooks/233-blip-visual-language-processing-with-output_files/233-blip-visual-language-processing-with-output_26_0.png rename to docs/notebooks/233-blip-visual-language-processing-with-output_files/233-blip-visual-language-processing-with-output_30_0.png diff --git a/docs/notebooks/233-blip-visual-language-processing-with-output_files/233-blip-visual-language-processing-with-output_7_0.png b/docs/notebooks/233-blip-visual-language-processing-with-output_files/233-blip-visual-language-processing-with-output_8_0.png similarity index 100% rename from docs/notebooks/233-blip-visual-language-processing-with-output_files/233-blip-visual-language-processing-with-output_7_0.png rename to docs/notebooks/233-blip-visual-language-processing-with-output_files/233-blip-visual-language-processing-with-output_8_0.png diff --git a/docs/notebooks/233-blip-visual-language-processing-with-output_files/index.html b/docs/notebooks/233-blip-visual-language-processing-with-output_files/index.html index 538fe3964b1..10f201080e3 100644 --- a/docs/notebooks/233-blip-visual-language-processing-with-output_files/index.html +++ b/docs/notebooks/233-blip-visual-language-processing-with-output_files/index.html @@ -1,9 +1,9 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/233-blip-visual-language-processing-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/233-blip-visual-language-processing-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/233-blip-visual-language-processing-with-output_files/


../
-233-blip-visual-language-processing-with-output..> 12-Jul-2023 00:11              206940
-233-blip-visual-language-processing-with-output..> 12-Jul-2023 00:11              210551
-233-blip-visual-language-processing-with-output..> 12-Jul-2023 00:11              210551
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/233-blip-visual-language-processing-with-output_files/


../
+233-blip-visual-language-processing-with-output..> 16-Aug-2023 01:31              206940
+233-blip-visual-language-processing-with-output..> 16-Aug-2023 01:31              210551
+233-blip-visual-language-processing-with-output..> 16-Aug-2023 01:31              210551
 

diff --git a/docs/notebooks/234-encodec-audio-compression-with-output.rst b/docs/notebooks/234-encodec-audio-compression-with-output.rst index 1c79b216849..309214879cd 100644 --- a/docs/notebooks/234-encodec-audio-compression-with-output.rst +++ b/docs/notebooks/234-encodec-audio-compression-with-output.rst @@ -1,6 +1,8 @@ Audio compression with EnCodec and OpenVINO =========================================== +.. _top: + Compression is an important part of the Internet today because it enables people to easily share high-quality photos, listen to audio messages, stream their favorite shows, and so much more. Even when using @@ -26,8 +28,26 @@ and original `repo `__. image.png -Prerequisites -------------- +**Table of contents**: + +- `Prerequisites <#prerequisites>`__ +- `Instantiate audio compression pipeline <#instantiate-audio-compression-pipeline>`__ +- `Explore EnCodec pipeline <#explore-encodec-pipeline>`__ + + - `Preprocessing <#preprocessing>`__ + - `Encoding <#encoding>`__ + - `Decompression <#decompression>`__ + +- `Convert model to OpenVINO Intermediate Representation format <#convert-model-to-openvino-intermediate-representation-format>`__ +- `Integrate OpenVINO to EnCodec pipeline <#integrate-openvino-to-encodec-pipeline>`__ + + - `Select inference device <#select-inference-device>`__ + +- `Run EnCodec with OpenVINO <#run-encodec-with-openvino>`__ + +Prerequisites `⇑ <#top>`__ +############################################################################################################################### + Install required dependencies: @@ -35,8 +55,9 @@ Install required dependencies: !python -W ignore -m pip install -q -r requirements.txt -Instantiate audio compression pipeline --------------------------------------- +Instantiate audio compression pipeline `⇑ <#top>`__ +############################################################################################################################### + `Codecs `__, which act as encoders and decoders for streams of data, help empower most of the audio @@ -74,9 +95,9 @@ stereophonic audio trained on music-only data. In this tutorial, we will use ``encodec_model_24khz`` as an example, but the same actions are also applicable to ``encodec_model_48khz`` model as well. To start working with this model, we need to instantiate model -class using EncodecModel.encodec_model_24khz() and select required +class using ``EncodecModel.encodec_model_24khz()`` and select required compression bandwidth among available: 1.5, 3, 6, 12 or 24 kbps for 24 -kHz model and 3, 6, 12 and 24 kbps for 48 kHz model. We will use 6 kbs +kHz model and 3, 6, 12 and 24 kbps for 48 kHz model. We will use 6 kbps bandwidth. .. |encodec_compression| image:: https://github.com/openvinotoolkit/openvino_notebooks/assets/29454499/5cd9a482-b42b-4dea-85a5-6d66b20ce13d @@ -93,8 +114,9 @@ bandwidth. model = EncodecModel.encodec_model_24khz() model.set_target_bandwidth(6.0) -Explore EnCodec pipeline ------------------------- +Explore EnCodec pipeline `⇑ <#top>`__ +############################################################################################################################### + Let us explore model capabilities on example audio: @@ -144,8 +166,9 @@ Let us explore model capabilities on example audio: .. image:: 234-encodec-audio-compression-with-output_files/234-encodec-audio-compression-with-output_6_2.png -Preprocessing -~~~~~~~~~~~~~ +Preprocessing `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + To achieve the best result, audio should have the number of channels and sample rate expected by the model. If audio does not fulfill these @@ -172,8 +195,9 @@ number of channels using the ``convert_audio`` function. wav = convert_audio(wav, sr, model_sr, model_channels) -Encoding -~~~~~~~~ +Encoding `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Audio waveform should be split by chunks and then encoded by Encoder model, then compressed by quantizer for reducing memory. The result of @@ -221,8 +245,9 @@ Let us compare obtained compression result: Great! Now, we see the power of hyper compression. Binary size of a file becomes 60 times smaller and more suitable for sending via network. -Decompression -~~~~~~~~~~~~~ +Decompression `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + After successful sending of the compressed audio, it should be decompressed on the recipient’s side. The decoder model is responsible @@ -270,8 +295,8 @@ audio. Nice! Audio sounds close to original. -Convert model to OpenVINO Intermediate Representation format ------------------------------------------------------------- +Convert model to OpenVINO Intermediate Representation format. `⇑ <#top>`__ +############################################################################################################################### For best results with OpenVINO, it is recommended to convert the model to OpenVINO IR format. OpenVINO supports PyTorch via ONNX conversion. We @@ -329,25 +354,25 @@ with ``openvino.runtime.serialize``. .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/encodec/modules/conv.py:60: TracerWarning: Converting a tensor to a Python float might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/encodec/modules/conv.py:60: TracerWarning: Converting a tensor to a Python float might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! ideal_length = (math.ceil(n_frames) - 1) * stride + (kernel_size - padding_total) - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/encodec/modules/conv.py:85: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/encodec/modules/conv.py:85: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! assert padding_left >= 0 and padding_right >= 0, (padding_left, padding_right) - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/encodec/modules/conv.py:87: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/encodec/modules/conv.py:87: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! max_pad = max(padding_left, padding_right) - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/encodec/modules/conv.py:89: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/encodec/modules/conv.py:89: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if length <= max_pad: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/symbolic_opset9.py:4315: UserWarning: Exporting a model to ONNX with a batch_size other than 1, with a variable length with LSTM can cause an error when running the ONNX model with a different batch size. Make sure to save the model with a batch size of 1, or define the initial states (h0/c0) as inputs of the model. + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/symbolic_opset9.py:4315: UserWarning: Exporting a model to ONNX with a batch_size other than 1, with a variable length with LSTM can cause an error when running the ONNX model with a different batch size. Make sure to save the model with a batch size of 1, or define the initial states (h0/c0) as inputs of the model. warnings.warn( - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/_internal/jit_utils.py:258: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/_internal/jit_utils.py:258: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) _C._jit_pass_onnx_node_shape_type_inference(node, params_dict, opset_version) - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) _C._jit_pass_onnx_graph_shape_type_inference( - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) _C._jit_pass_onnx_graph_shape_type_inference( - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) _C._jit_pass_onnx_graph_shape_type_inference( - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) _C._jit_pass_onnx_graph_shape_type_inference( @@ -364,16 +389,17 @@ with ``openvino.runtime.serialize``. .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/encodec/quantization/core_vq.py:358: TracerWarning: torch.tensor results are registered as constants in the trace. You can safely ignore this warning if you use this function to create tensors out of constant variables that would be the same every time you call this function. In any other case, this might cause the trace to be incorrect. + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/encodec/quantization/core_vq.py:358: TracerWarning: torch.tensor results are registered as constants in the trace. You can safely ignore this warning if you use this function to create tensors out of constant variables that would be the same every time you call this function. In any other case, this might cause the trace to be incorrect. quantized_out = torch.tensor(0.0, device=q_indices.device) - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/encodec/quantization/core_vq.py:359: TracerWarning: Iterating over a tensor might cause the trace to be incorrect. Passing a tensor of different shape won't change the number of iterations executed (and might lead to errors or silently give incorrect results). + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/encodec/quantization/core_vq.py:359: TracerWarning: Iterating over a tensor might cause the trace to be incorrect. Passing a tensor of different shape won't change the number of iterations executed (and might lead to errors or silently give incorrect results). for i, indices in enumerate(q_indices): - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/encodec/modules/conv.py:103: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/encodec/modules/conv.py:103: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! assert (padding_left + padding_right) <= x.shape[-1] -Integrate OpenVINO to EnCodec pipeline --------------------------------------- +Integrate OpenVINO to EnCodec pipeline `⇑ <#top>`__ +############################################################################################################################### + The following steps are required for integration of OpenVINO to EnCodec pipeline: @@ -383,14 +409,40 @@ pipeline: 3. Replace the original frame processing functions with OpenVINO based algorithms. +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + .. code:: ipython3 - device = "CPU" + import ipywidgets as widgets - compiled_encoder = core.compile_model(encoder_ov, device) + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + compiled_encoder = core.compile_model(encoder_ov, device.value) encoder_out = compiled_encoder.output(0) - compiled_decoder = core.compile_model(decoder_ov, device) + compiled_decoder = core.compile_model(decoder_ov, device.value) decoder_out = compiled_decoder.output(0) .. code:: ipython3 @@ -422,8 +474,9 @@ pipeline: model._encode_frame = encode_frame model._decode_frame = decode_frame -Run EnCodec with OpenVINO -------------------------- +Run EnCodec with OpenVINO `⇑ <#top>`__ +############################################################################################################################### + The process of running encodec with OpenVINO under hood will be the same like with the original PyTorch models. @@ -475,7 +528,7 @@ like with the original PyTorch models. -.. image:: 234-encodec-audio-compression-with-output_files/234-encodec-audio-compression-with-output_36_1.png +.. image:: 234-encodec-audio-compression-with-output_files/234-encodec-audio-compression-with-output_38_1.png .. code:: ipython3 diff --git a/docs/notebooks/234-encodec-audio-compression-with-output_files/234-encodec-audio-compression-with-output_36_1.png b/docs/notebooks/234-encodec-audio-compression-with-output_files/234-encodec-audio-compression-with-output_38_1.png similarity index 100% rename from docs/notebooks/234-encodec-audio-compression-with-output_files/234-encodec-audio-compression-with-output_36_1.png rename to docs/notebooks/234-encodec-audio-compression-with-output_files/234-encodec-audio-compression-with-output_38_1.png diff --git a/docs/notebooks/234-encodec-audio-compression-with-output_files/index.html b/docs/notebooks/234-encodec-audio-compression-with-output_files/index.html index b6d8d7fc475..a45a6a0a7c7 100644 --- a/docs/notebooks/234-encodec-audio-compression-with-output_files/index.html +++ b/docs/notebooks/234-encodec-audio-compression-with-output_files/index.html @@ -1,9 +1,9 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/234-encodec-audio-compression-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/234-encodec-audio-compression-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/234-encodec-audio-compression-with-output_files/


../
-234-encodec-audio-compression-with-output_19_1.png 12-Jul-2023 00:11               44358
-234-encodec-audio-compression-with-output_36_1.png 12-Jul-2023 00:11               44009
-234-encodec-audio-compression-with-output_6_2.png  12-Jul-2023 00:11               45005
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/234-encodec-audio-compression-with-output_files/


../
+234-encodec-audio-compression-with-output_19_1.png 16-Aug-2023 01:31               44358
+234-encodec-audio-compression-with-output_38_1.png 16-Aug-2023 01:31               44009
+234-encodec-audio-compression-with-output_6_2.png  16-Aug-2023 01:31               45005
 

diff --git a/docs/notebooks/235-controlnet-stable-diffusion-with-output.rst b/docs/notebooks/235-controlnet-stable-diffusion-with-output.rst index 087ed94362c..1ce9e215d76 100644 --- a/docs/notebooks/235-controlnet-stable-diffusion-with-output.rst +++ b/docs/notebooks/235-controlnet-stable-diffusion-with-output.rst @@ -1,6 +1,8 @@ Text-to-Image Generation with ControlNet Conditioning ===================================================== +.. _top: + Diffusion models make a revolution in AI-generated art. This technology enables creation of high-quality images simply by writing a text prompt. Even though this technology gives very promising results, the diffusion @@ -51,11 +53,14 @@ This is the key difference between standard diffusion and latent diffusion models: in latent diffusion, the model is trained to generate latent (compressed) representations of the images. -There are three main components in latent diffusion: +There are three main components in latent diffusion: -* A text-encoder, for example `CLIP’s Text Encoder `__ for creation condition to generate image from text prompt. -* A U-Net for step-by-step denoising latent image representation. -* An autoencoder (VAE) for encoding input image to latent space (if required) and decoding latent space to image back after generation. +- A text-encoder, for example `CLIP’s Text + Encoder `__ + for creation condition to generate image from text prompt. +- A U-Net for step-by-step denoising latent image representation. +- An autoencoder (VAE) for encoding input image to latent space (if + required) and decoding latent space to image back after generation. For more details regarding Stable Diffusion work, refer to the `project website `__. @@ -105,7 +110,7 @@ In the end, we are left with a very similar image synthesis pipeline with an additional control added for the shape of the output features in the final image. -Training ControlNet comprises of the following steps: +Training ControlNet consists of the following steps: 1. Cloning the pre-trained parameters of a Diffusion model, such as Stable Diffusion’s latent UNet, (referred to as “trainable copy”) @@ -136,26 +141,52 @@ of the target in the image: This tutorial focuses mainly on conditioning by pose. However, the discussed steps are also applicable to other annotation modes. -Prerequisites -------------- +**Table of contents**: + +- `Prerequisites <#prerequisites>`__ +- `Instantiating Generation Pipeline <#instantiating-generation-pipeline>`__ + + - `ControlNet in Diffusers library <#controlnet-in-diffusers-library>`__ + - `OpenPose <#openpose>`__ + +- `Convert models to OpenVINO Intermediate representation (IR) format <#convert-models-to-openvino-intermediate-representation-ir-format>`__ + + - `OpenPose conversion <#openpose-conversion>`__ + +- `Select inference device <#select-inference-device>`__ + + - `ControlNet conversion <#controlnet-conversion>`__ + - `UNet conversion <#unet-conversion>`__ + - `Text Encoder <#text-encoder>`__ + - `VAE Decoder conversion <#vae-decoder-conversion>`__ + +- `Prepare Inference pipeline <#prepare-inference-pipeline>`__ +- `Running Text-to-Image Generation with ControlNet Conditioning and OpenVINO <#running-text-to-image-generation-with-controlnet-conditioning-and-openvino>`__ +- `Select inference device <#select-inference-device>`__ + +Prerequisites `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 - !pip install -q "diffusers>=0.14.0" "git+https://github.com/huggingface/accelerate.git" controlnet-aux gradio + !pip install -q "diffusers==0.14.0" "controlnet-aux>=0.0.6" "gradio>=3.36" .. parsed-literal:: - [notice] A new release of pip available: 22.3.1 -> 23.0.1 + [notice] A new release of pip is available: 23.1.2 -> 23.2 [notice] To update, run: pip install --upgrade pip -Instantiating Generation Pipeline ---------------------------------- +Instantiating Generation Pipeline `⇑ <#top>`__ +############################################################################################################################### + + +ControlNet in Diffusers library `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -ControlNet in Diffusers library -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ For working with Stable Diffusion and ControlNet models, we will use Hugging Face `Diffusers `__ @@ -176,7 +207,7 @@ controlnet model and ``stable-diffusion-v1-5``: import torch from diffusers import StableDiffusionControlNetPipeline, ControlNetModel - controlnet = ControlNetModel.from_pretrained("lllyasviel/sd-controlnet-openpose", torch_dtype=torch.float32) + controlnet = ControlNetModel.from_pretrained("lllyasviel/control_v11p_sd15_openpose", torch_dtype=torch.float32) pipe = StableDiffusionControlNetPipeline.from_pretrained( "runwayml/stable-diffusion-v1-5", controlnet=controlnet ) @@ -184,23 +215,16 @@ controlnet model and ``stable-diffusion-v1-5``: .. parsed-literal:: - 2023-03-12 15:26:19.533980: I tensorflow/core/util/util.cc:169] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-07-16 15:33:13.040077: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-07-16 15:33:13.079142: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. + 2023-07-16 15:33:13.688517: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + `text_config_dict` is provided which will be used to initialize `CLIPTextConfig`. The value `text_config["id2label"]` will be overriden. +OpenPose `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -.. parsed-literal:: - - Fetching 15 files: 0%| | 0/15 [00:00`__ @@ -225,6 +249,13 @@ The code below demonstrates how to instantiate the OpenPose model. pose_estimator = OpenposeDetector.from_pretrained("lllyasviel/ControlNet") + +.. parsed-literal:: + + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/controlnet_aux/mediapipe_face/mediapipe_face_common.py:7: UserWarning: The module 'mediapipe' is not installed. The package will have limited functionality. Please install it using the command: pip install 'mediapipe' + warnings.warn( + + Now, let us check its result on example image: .. code:: ipython3 @@ -281,8 +312,8 @@ Now, let us check its result on example image: .. image:: 235-controlnet-stable-diffusion-with-output_files/235-controlnet-stable-diffusion-with-output_8_0.png -Convert models to OpenVINO Intermediate representation (IR) format ------------------------------------------------------------------- +Convert models to OpenVINO Intermediate representation (IR) format. `⇑ <#top>`__ +############################################################################################################################### OpenVINO supports PyTorch through export to the ONNX format. We will use the ``torch.onnx.export`` function for obtaining the ONNX model, we can @@ -295,23 +326,25 @@ example, input and output names or dynamic shapes). While ONNX models are directly supported by OpenVINO™ runtime, it can be useful to convert them to IR format to take the advantage of advanced -OpenVINO optimization tools and features. We will use OpenVINO `Model -Optimizer `__ +OpenVINO optimization tools and features. We will use `model conversion +API `__ to convert a model to IR format and compression weights to ``FP16`` format. The pipeline consists of five important parts: -* OpenPose for obtaining annotation based on an estimated pose. -* ControlNet for conditioning by image annotation. -* Text Encoder for creation condition to generate an image from a text prompt. -* Unet for step-by-step denoising latent image representation. -* Autoencoder (VAE) for decoding latent space to image. +- OpenPose for obtaining annotation based on an estimated pose. +- ControlNet for conditioning by image annotation. +- Text Encoder for creation condition to generate an image from a text + prompt. +- Unet for step-by-step denoising latent image representation. +- Autoencoder (VAE) for decoding latent space to image. Let us convert each part: -OpenPose conversion -~~~~~~~~~~~~~~~~~~~ +OpenPose conversion `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + OpenPose model is represented in the pipeline as a wrapper on the PyTorch model which not only detects poses on an input image but is also @@ -348,6 +381,7 @@ model with the OpenVINO model, using the following code: .. code:: ipython3 from openvino.runtime import Model, Core + from collections import namedtuple class OpenPoseOVModel: @@ -385,10 +419,46 @@ model with the OpenVINO model, using the following code: """ self.model.reshape({0: [1, 3, height, width]}) self.compiled_model = self.core.compile_model(self.model) + + def parameters(self): + Device = namedtuple("Device", ["device"]) + return [Device(torch.device("cpu"))] + core = Core() - ov_openpose = OpenPoseOVModel(core, OPENPOSE_OV_PATH) + +Select inference device `⇑ <#top>`__ +############################################################################################################################### + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=2, options=('CPU', 'GPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + ov_openpose = OpenPoseOVModel(core, OPENPOSE_OV_PATH, device=device.value) pose_estimator.body_estimation.model = ov_openpose .. code:: ipython3 @@ -398,23 +468,27 @@ model with the OpenVINO model, using the following code: -.. image:: 235-controlnet-stable-diffusion-with-output_files/235-controlnet-stable-diffusion-with-output_14_0.png +.. image:: 235-controlnet-stable-diffusion-with-output_files/235-controlnet-stable-diffusion-with-output_17_0.png Great! As we can see, it works perfectly. -ControlNet conversion -~~~~~~~~~~~~~~~~~~~~~ +ControlNet conversion `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -The controlNet model accepts the same inputs like UNet in Stable + +The ControlNet model accepts the same inputs like UNet in Stable Diffusion pipeline and additional condition sample - skeleton key points map predicted by pose estimator: -* ``sample`` - latent image sample from the previous step, generation process has not been started yet, so we will use random noise, -* ``timestep`` - current scheduler step, -* ``encoder_hidden_state`` - hidden state of text encoder, -* ``controlnet_cond`` - condition input annotation. The output of the model is attention hidden states from down and middle blocks, which serves additional context for the UNet model. +- ``sample`` - latent image sample from the previous step, generation + process has not been started yet, so we will use random noise, +- ``timestep`` - current scheduler step, +- ``encoder_hidden_state`` - hidden state of text encoder, +- ``controlnet_cond`` - condition input annotation. +The output of the model is attention hidden states from down and middle +blocks, which serves additional context for the UNet model. .. code:: ipython3 @@ -455,8 +529,9 @@ map predicted by pose estimator: ControlNet will be loaded from controlnet-pose.xml -UNet conversion -~~~~~~~~~~~~~~~ +UNet conversion `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The process of UNet model conversion remains the same, like for original Stable Diffusion model, but with respect to the new inputs generated by @@ -501,18 +576,17 @@ ControlNet. .. parsed-literal:: - 4989 + 5513 -Text Encoder -~~~~~~~~~~~~ +Text Encoder `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -The text-encoder is responsible for transforming the input prompt, for -example, “a photo of an astronaut riding a horse” into an embedding -space that can be understood by the U-Net. It is usually a simple -transformer-based encoder that maps a sequence of input tokens to a -sequence of latent text embeddings. +The text-encoder is responsible for transforming the input prompt, for example, +“a photo of an astronaut riding a horse” into an embedding space that can be +understood by the U-Net. It is usually a simple transformer-based encoder that +maps a sequence of input tokens to a sequence of latent text embeddings. The input of the text encoder is tensor ``input_ids``, which contains indexes of tokens from text processed by the tokenizer and padded to the @@ -584,8 +658,9 @@ this opset. -VAE Decoder conversion -~~~~~~~~~~~~~~~~~~~~~~ +VAE Decoder conversion `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The VAE model has two parts, an encoder, and a decoder. The encoder is used to convert the image into a low-dimensional latent representation, @@ -651,8 +726,9 @@ diffusion VAE decoder will be loaded from vae_decoder.xml -Prepare Inference pipeline --------------------------- +Prepare Inference pipeline `⇑ <#top>`__ +############################################################################################################################### + Putting it all together, let us now take a closer look at how the model works in inference by illustrating the logical flow. |detailed workflow| @@ -1053,6 +1129,13 @@ on OpenVINO. image = np.transpose(image, (0, 2, 3, 1)) return image + +.. parsed-literal:: + + /tmp/ipykernel_1180132/670611772.py:1: FutureWarning: Importing `DiffusionPipeline` or `ImagePipelineOutput` from diffusers.pipeline_utils is deprecated. Please import from diffusers.pipelines.pipeline_utils instead. + from diffusers.pipeline_utils import DiffusionPipeline + + .. code:: ipython3 from transformers import CLIPTokenizer @@ -1099,8 +1182,8 @@ on OpenVINO. fig.savefig("result.png", bbox_inches='tight') return fig -Running Text-to-Image Generation with ControlNet Conditioning and OpenVINO --------------------------------------------------------------------------- +Running Text-to-Image Generation with ControlNet Conditioning and OpenVINO. `⇑ <#top>`__ +############################################################################################################################### Now, we are ready to start generation. For improving the generation process, we also introduce an opportunity to provide a @@ -1112,9 +1195,37 @@ this We can keep this field empty if we want to generate image without negative prompting. +Select inference device `⇑ <#top>`__ +############################################################################################################################### + + +Select device from dropdown list for running inference using OpenVINO: + .. code:: ipython3 - ov_pipe = OVContrlNetStableDiffusionPipeline(tokenizer, scheduler, core, CONTROLNET_OV_PATH, TEXT_ENCODER_OV_PATH, UNET_OV_PATH, VAE_DECODER_OV_PATH, device="AUTO") + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='CPU', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', options=('CPU', 'GPU', 'AUTO'), value='CPU') + + + +.. code:: ipython3 + + ov_pipe = OVContrlNetStableDiffusionPipeline(tokenizer, scheduler, core, CONTROLNET_OV_PATH, TEXT_ENCODER_OV_PATH, UNET_OV_PATH, VAE_DECODER_OV_PATH, device=device.value) .. code:: ipython3 @@ -1157,4 +1268,19 @@ negative prompting. pose_btn.click(extract_pose, inp_img, [out_pose, step1, step2]) btn.click(generate, [out_pose, inp_prompt, inp_neg_prompt, inp_seed, inp_steps], out_result) - demo.queue().launch(share=True, debug=True) + demo.queue().launch(share=True) + + +.. parsed-literal:: + + Running on local URL: http://127.0.0.1:7860 + Running on public URL: https://6927b0a05729fd4297.gradio.live + + This share link expires in 72 hours. For free permanent hosting and GPU upgrades, run `gradio deploy` from Terminal to deploy to Spaces (https://huggingface.co/spaces) + + + +.. raw:: html + +
+ diff --git a/docs/notebooks/235-controlnet-stable-diffusion-with-output_files/235-controlnet-stable-diffusion-with-output_14_0.png b/docs/notebooks/235-controlnet-stable-diffusion-with-output_files/235-controlnet-stable-diffusion-with-output_14_0.png deleted file mode 100644 index 7188332bf4f..00000000000 --- a/docs/notebooks/235-controlnet-stable-diffusion-with-output_files/235-controlnet-stable-diffusion-with-output_14_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:7a81e4d1d0073f442685f6aa5aa9570e02a572bc0ecf71afc39c8a77ab8b8e8d -size 492236 diff --git a/docs/notebooks/235-controlnet-stable-diffusion-with-output_files/235-controlnet-stable-diffusion-with-output_17_0.png b/docs/notebooks/235-controlnet-stable-diffusion-with-output_files/235-controlnet-stable-diffusion-with-output_17_0.png new file mode 100644 index 00000000000..7af8840ee7a --- /dev/null +++ b/docs/notebooks/235-controlnet-stable-diffusion-with-output_files/235-controlnet-stable-diffusion-with-output_17_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:db732e96aa0954fadfbe1bdd5cddd2131ca83d526ae211d1da75625365b1a482 +size 498463 diff --git a/docs/notebooks/235-controlnet-stable-diffusion-with-output_files/235-controlnet-stable-diffusion-with-output_8_0.png b/docs/notebooks/235-controlnet-stable-diffusion-with-output_files/235-controlnet-stable-diffusion-with-output_8_0.png index 26b2284ee63..7af8840ee7a 100644 --- a/docs/notebooks/235-controlnet-stable-diffusion-with-output_files/235-controlnet-stable-diffusion-with-output_8_0.png +++ b/docs/notebooks/235-controlnet-stable-diffusion-with-output_files/235-controlnet-stable-diffusion-with-output_8_0.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:045843cd549bd8a56e91319a52766bf17ebc63d0d0b4b73313b472ec0192b53e -size 491726 +oid sha256:db732e96aa0954fadfbe1bdd5cddd2131ca83d526ae211d1da75625365b1a482 +size 498463 diff --git a/docs/notebooks/235-controlnet-stable-diffusion-with-output_files/index.html b/docs/notebooks/235-controlnet-stable-diffusion-with-output_files/index.html index 1f98f7b386a..631e49636d2 100644 --- a/docs/notebooks/235-controlnet-stable-diffusion-with-output_files/index.html +++ b/docs/notebooks/235-controlnet-stable-diffusion-with-output_files/index.html @@ -1,8 +1,8 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/235-controlnet-stable-diffusion-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/235-controlnet-stable-diffusion-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/235-controlnet-stable-diffusion-with-output_files/


../
-235-controlnet-stable-diffusion-with-output_14_..> 12-Jul-2023 00:11              492236
-235-controlnet-stable-diffusion-with-output_8_0..> 12-Jul-2023 00:11              491726
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/235-controlnet-stable-diffusion-with-output_files/


../
+235-controlnet-stable-diffusion-with-output_17_..> 16-Aug-2023 01:31              498463
+235-controlnet-stable-diffusion-with-output_8_0..> 16-Aug-2023 01:31              498463
 

diff --git a/docs/notebooks/236-stable-diffusion-v2-infinite-zoom-with-output.rst b/docs/notebooks/236-stable-diffusion-v2-infinite-zoom-with-output.rst index 4d3c7ac858a..a07b5c22ebb 100644 --- a/docs/notebooks/236-stable-diffusion-v2-infinite-zoom-with-output.rst +++ b/docs/notebooks/236-stable-diffusion-v2-infinite-zoom-with-output.rst @@ -1,6 +1,8 @@ Infinite Zoom Stable Diffusion v2 and OpenVINO™ =============================================== +.. _top: + Stable Diffusion v2 is the next generation of Stable Diffusion model a Text-to-Image latent diffusion model created by the researchers and engineers from `Stability AI `__ and @@ -31,16 +33,35 @@ Stable Diffusion v2: What’s new? The new stable diffusion model offers a bunch of new features inspired by the other models that have emerged since the introduction of the first iteration. Some of the features that can be found in the new model -are: +are: -* The model comes with a new robust encoder, OpenCLIP, created by LAION and aided by Stability AI; this version v2 significantly enhances the produced photos over the V1 versions. -* The model can now generate images in a 768x768 resolution, offering more information to be shown in the generated images. -* The model finetuned with `v-objective `__. The v-parameterization is particularly useful for numerical stability throughout the diffusion process to enable progressive distillation for models. For models that operate at higher resolution, it is also discovered that the v-parameterization avoids color shifting artifacts that are known to affect high resolution diffusion models, and in the video setting it avoids temporal color shifting that sometimes appears with epsilon-prediction used in Stable Diffusion v1. -* The model also comes with a new diffusion model capable of running upscaling on the images generated. Upscaled images can be adjusted up to 4 times the original image. Provided as separated model, for more details please check `stable-diffusion-x4-upscaler `__ -* The model comes with a new refined depth architecture capable of preserving context from prior generation layers in an img2img setting. -This structure preservation helps generate images that preserving forms -and shadow of objects, but with different content. -* The model comes with an updated inpainting module built upon the previous model. This text-guided inpainting makes switching out parts in the image easier than before. +- The model comes with a new robust encoder, OpenCLIP, created by LAION + and aided by Stability AI; this version v2 significantly enhances the + produced photos over the V1 versions. +- The model can now generate images in a 768x768 resolution, offering + more information to be shown in the generated images. +- The model finetuned with + `v-objective `__. The + v-parameterization is particularly useful for numerical stability + throughout the diffusion process to enable progressive distillation + for models. For models that operate at higher resolution, it is also + discovered that the v-parameterization avoids color shifting + artifacts that are known to affect high resolution diffusion models, + and in the video setting it avoids temporal color shifting that + sometimes appears with epsilon-prediction used in Stable Diffusion + v1. +- The model also comes with a new diffusion model capable of running + upscaling on the images generated. Upscaled images can be adjusted up + to 4 times the original image. Provided as separated model, for more + details please check + `stable-diffusion-x4-upscaler `__ +- The model comes with a new refined depth architecture capable of + preserving context from prior generation layers in an image-to-image + setting. This structure preservation helps generate images that + preserving forms and shadow of objects, but with different content. +- The model comes with an updated inpainting module built upon the + previous model. This text-guided inpainting makes switching out parts + in the image easier than before. This notebook demonstrates how to convert and run Stable Diffusion v2 model using OpenVINO. @@ -48,11 +69,30 @@ model using OpenVINO. Notebook contains the following steps: 1. Convert PyTorch models to ONNX format. -2. Convert ONNX models to OpenVINO IR format, using Model Optimizer tool. -3. Run Stable Diffusion v2 inpainting pipeline for generation infinity zoom video. +2. Convert ONNX models to OpenVINO IR format, using model conversion + API. +3. Run Stable Diffusion v2 inpainting pipeline for generation infinity + zoom video + +**Table of contents**: + +- `Stable Diffusion v2 Infinite Zoom Showcase <#stable-diffusion-v2-infinite-zoom-showcase>`__ + + - `Stable Diffusion Text guided Inpainting <#stable-diffusion-text-guided-inpainting>`__ + +- `Prerequisites <#prerequisites>`__ + + - `Stable Diffusion in Diffusers library <#stable-diffusion-in-diffusers-library>`__ + - `Convert models to OpenVINO Intermediate representation (IR) format <#convert-models-to-openvino-intermediate-representation-ir-format>`__ + - `Prepare Inference pipeline <#prepare-inference-pipeline>`__ + - `Zoom Video Generation <#zoom-video-generation>`__ + - `Configure Inference Pipeline <#configure-inference-pipeline>`__ + - `Select inference device <#select-inference-device>`__ + - `Run Infinite Zoom video generation <#run-infinite-zoom-video-generation>`__ + +Stable Diffusion v2 Infinite Zoom Showcase `⇑ <#top>`__ +############################################################################################################################### -Stable Diffusion v2 Infinite Zoom Showcase ------------------------------------------- In this tutorial we consider how to use Stable Diffusion v2 model for generation sequence of images for infinite zoom video effect. To do @@ -60,13 +100,12 @@ this, we will need `stabilityai/stable-diffusion-2-inpainting `__ model. -Stable Diffusion Text guided Inpainting -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Stable Diffusion Text guided Inpainting `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -In image editing, inpainting is a process of restoring missing parts of -pictures. Most commonly applied to reconstructing old deteriorated -images, removing cracks, scratches, dust spots, or red-eyes from -photographs. +In image editing, inpainting is a process of restoring missing parts of pictures. Most +commonly applied to reconstructing old deteriorated images, removing +cracks, scratches, dust spots, or red-eyes from photographs. But with the power of AI and the Stable Diffusion model, inpainting can be used to achieve more than that. For example, instead of just @@ -74,7 +113,7 @@ restoring missing parts of an image, it can be used to render something entirely new in any part of an existing picture. Only your imagination limits it. -The workflow diagram explains how Stable Diffusion inpaining pipeline +The workflow diagram explains how Stable Diffusion inpainting pipeline for inpainting works: .. figure:: https://github.com/openvinotoolkit/openvino_notebooks/assets/22090501/9ac6de45-186f-4a3c-aa20-825825a337eb @@ -94,10 +133,10 @@ Using this inpainting feature, decreasing image by certain margin and masking this border for every new frame we can create interesting Zoom Out video based on our prompt. -Prerequisites -------------- +Prerequisites `⇑ <#top>`__ +############################################################################################################################### -install required packages +Install required packages: .. code:: ipython3 @@ -107,12 +146,12 @@ install required packages .. parsed-literal:: - [notice] A new release of pip available: 22.3.1 -> 23.0.1 + [notice] A new release of pip is available: 23.1.2 -> 23.2 [notice] To update, run: pip install --upgrade pip -Stable Diffusion in Diffusers library -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Stable Diffusion in Diffusers library `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ To work with Stable Diffusion v2, we will use Hugging Face `Diffusers `__ library. To @@ -135,16 +174,96 @@ The code below demonstrates how to create scheduler_inpaint = DPMSolverMultistepScheduler.from_config(pipe_inpaint.scheduler.config) +.. parsed-literal:: + + 2023-07-16 15:45:16.540634: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-07-16 15:45:16.577870: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. + 2023-07-16 15:45:17.175991: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + + + +.. parsed-literal:: + + Downloading (…)ain/model_index.json: 0%| | 0.00/544 [00:00`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ Conversion part of model stayed remain as in `Text-to-Image generation notebook <./236-stable-diffusion-v2-text-to-image.ipynb>`__. Except @@ -329,26 +448,25 @@ generated latents channels + 4 for latent representation of masked image .. parsed-literal:: - /tmp/ipykernel_384919/3505677505.py:19: FutureWarning: 'torch.onnx._export' is deprecated in version 1.12.0 and will be removed in version 1.14. Please use `torch.onnx.export` instead. + /tmp/ipykernel_1181138/3505677505.py:19: FutureWarning: 'torch.onnx._export' is deprecated in version 1.12.0 and will be removed in version 1.14. Please use `torch.onnx.export` instead. torch.onnx._export( - /home/ea/work/transformers/src/transformers/models/clip/modeling_clip.py:759: TracerWarning: torch.tensor results are registered as constants in the trace. You can safely ignore this warning if you use this function to create tensors out of constant variables that would be the same every time you call this function. In any other case, this might cause the trace to be incorrect. - mask.fill_(torch.tensor(torch.finfo(dtype).min)) - /home/ea/work/transformers/src/transformers/models/clip/modeling_clip.py:284: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/transformers/models/clip/modeling_clip.py:684: TracerWarning: torch.tensor results are registered as constants in the trace. You can safely ignore this warning if you use this function to create tensors out of constant variables that would be the same every time you call this function. In any other case, this might cause the trace to be incorrect. + mask = torch.full((tgt_len, tgt_len), torch.tensor(torch.finfo(dtype).min, device=device), device=device) + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/transformers/models/clip/modeling_clip.py:284: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if attn_weights.size() != (bsz * self.num_heads, tgt_len, src_len): - /home/ea/work/transformers/src/transformers/models/clip/modeling_clip.py:292: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/transformers/models/clip/modeling_clip.py:292: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if causal_attention_mask.size() != (bsz, 1, tgt_len, src_len): - /home/ea/work/transformers/src/transformers/models/clip/modeling_clip.py:324: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/transformers/models/clip/modeling_clip.py:324: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if attn_output.size() != (bsz * self.num_heads, tgt_len, self.head_dim): - /home/ea/work/notebooks_env/lib/python3.8/site-packages/torch/onnx/symbolic_helper.py:710: UserWarning: Type cannot be inferred, which might cause exported graph to produce incorrect results. + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/torch/onnx/symbolic_helper.py:710: UserWarning: Type cannot be inferred, which might cause exported graph to produce incorrect results. warnings.warn( - /home/ea/work/notebooks_env/lib/python3.8/site-packages/torch/onnx/symbolic_opset9.py:5408: UserWarning: Exporting aten::index operator of advanced indexing in opset 14 is achieved by combination of multiple ONNX operators, including Reshape, Transpose, Concat, and Gather. If indices include negative values, the exported graph will produce incorrect results. + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/torch/onnx/symbolic_opset9.py:5408: UserWarning: Exporting aten::index operator of advanced indexing in opset 14 is achieved by combination of multiple ONNX operators, including Reshape, Transpose, Concat, and Gather. If indices include negative values, the exported graph will produce incorrect results. warnings.warn( .. parsed-literal:: Text Encoder successfully converted to ONNX - Warning: One or more of the values of the Constant can't fit in the float16 data type. Those values were casted to the nearest limit value, the model can produce incorrect results. [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html [ SUCCESS ] Generated IR version 11 model. @@ -375,19 +493,19 @@ generated latents channels + 4 for latent representation of masked image .. parsed-literal:: - /tmp/ipykernel_384919/3505677505.py:56: FutureWarning: 'torch.onnx._export' is deprecated in version 1.12.0 and will be removed in version 1.14. Please use `torch.onnx.export` instead. + /tmp/ipykernel_1181138/3505677505.py:56: FutureWarning: 'torch.onnx._export' is deprecated in version 1.12.0 and will be removed in version 1.14. Please use `torch.onnx.export` instead. torch.onnx._export( - /home/ea/work/diffusers/src/diffusers/models/unet_2d_condition.py:526: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/diffusers/models/unet_2d_condition.py:752: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if any(s % default_overall_up_factor != 0 for s in sample.shape[-2:]): - /home/ea/work/diffusers/src/diffusers/models/resnet.py:185: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/diffusers/models/resnet.py:214: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! assert hidden_states.shape[1] == self.channels - /home/ea/work/diffusers/src/diffusers/models/resnet.py:190: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/diffusers/models/resnet.py:219: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! assert hidden_states.shape[1] == self.channels - /home/ea/work/diffusers/src/diffusers/models/resnet.py:112: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/diffusers/models/resnet.py:138: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! assert hidden_states.shape[1] == self.channels - /home/ea/work/diffusers/src/diffusers/models/resnet.py:125: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/diffusers/models/resnet.py:151: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if hidden_states.shape[0] >= 64: - /home/ea/work/diffusers/src/diffusers/models/unet_2d_condition.py:651: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/diffusers/models/unet_2d_condition.py:977: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if not return_dict: @@ -429,11 +547,11 @@ generated latents channels + 4 for latent representation of masked image .. parsed-literal:: - /home/ea/work/notebooks_env/lib/python3.8/site-packages/torch/onnx/_internal/jit_utils.py:258: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/torch/onnx/_internal/jit_utils.py:258: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) _C._jit_pass_onnx_node_shape_type_inference(node, params_dict, opset_version) - /home/ea/work/notebooks_env/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) _C._jit_pass_onnx_graph_shape_type_inference( - /home/ea/work/notebooks_env/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) _C._jit_pass_onnx_graph_shape_type_inference( @@ -450,11 +568,11 @@ generated latents channels + 4 for latent representation of masked image .. parsed-literal:: - /home/ea/work/notebooks_env/lib/python3.8/site-packages/torch/onnx/_internal/jit_utils.py:258: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/torch/onnx/_internal/jit_utils.py:258: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) _C._jit_pass_onnx_node_shape_type_inference(node, params_dict, opset_version) - /home/ea/work/notebooks_env/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) _C._jit_pass_onnx_graph_shape_type_inference( - /home/ea/work/notebooks_env/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) _C._jit_pass_onnx_graph_shape_type_inference( @@ -469,10 +587,11 @@ generated latents channels + 4 for latent representation of masked image VAE decoder successfully converted to IR -Prepare Inference pipeline -~~~~~~~~~~~~~~~~~~~~~~~~~~ +Prepare Inference pipeline `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -As it was descussed previously, Inpainting inference pipeline is based + +As it was discussed previously, Inpainting inference pipeline is based on Text-to-Image inference pipeline with addition mask processing step. We will reuse ``OVStableDiffusionPipeline`` basic utilities in ``OVStableDiffusionInpaintingPipeline`` class. @@ -539,6 +658,13 @@ We will reuse ``OVStableDiffusionPipeline`` basic utilities in return mask, masked_image + +.. parsed-literal:: + + /tmp/ipykernel_1181138/859685649.py:8: FutureWarning: Importing `DiffusionPipeline` or `ImagePipelineOutput` from diffusers.pipeline_utils is deprecated. Please import from diffusers.pipelines.pipeline_utils instead. + from diffusers.pipeline_utils import DiffusionPipeline + + .. code:: ipython3 class OVStableDiffusionInpaintingPipeline(DiffusionPipeline): @@ -899,14 +1025,15 @@ We will reuse ``OVStableDiffusionPipeline`` basic utilities in return timesteps, num_inference_steps - t_start -Zoom Video Generation -~~~~~~~~~~~~~~~~~~~~~ +Zoom Video Generation `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + For achieving zoom effect, we will use inpainting to expand images beyond their original borders. We run our -OVStableDiffusionInpaintingPipeline in the loop, where each next frame -will add edges to previous. The frame generation process illustrated on -diagram below: +``OVStableDiffusionInpaintingPipeline`` in the loop, where each next +frame will add edges to previous. The frame generation process +illustrated on diagram below: .. figure:: https://user-images.githubusercontent.com/29454499/228739686-436f2759-4c79-42a2-a70f-959fb226834c.png :alt: frame generation) @@ -920,10 +1047,10 @@ image scaling. There are 2 zooming directions: -* Zoom Out - move away from object -* Zoom In - move closer to object +- Zoom Out - move away from object +- Zoom In - move closer to object -Zoom In will be processed in the same way like Zoom Out, but after +Zoom In will be processed in the same way as Zoom Out, but after generation is finished, we record frames in reversed order. .. code:: ipython3 @@ -1134,14 +1261,13 @@ generation is finished, we record frames in reversed order. loop=0, ) -Configure Inference Pipeline -~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Configure Inference Pipeline `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Configuration steps: -1. Load models on device. -2. Configure tokenizer and scheduler. -3. Create instance of OvStableDiffusionInpaintingPipeline class. +Configuration steps: 1. Load models on device 2. Configure tokenizer and +scheduler 3. Create instance of ``OVStableDiffusionInpaintingPipeline`` +class .. code:: ipython3 @@ -1150,11 +1276,42 @@ Configuration steps: core = Core() tokenizer = CLIPTokenizer.from_pretrained('openai/clip-vit-large-patch14') + +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets - text_enc_inpaint = core.compile_model(TEXT_ENCODER_OV_PATH_INPAINT, "CPU") - unet_model_inpaint = core.compile_model(UNET_OV_PATH_INPAINT, "CPU") - vae_decoder_inpaint = core.compile_model(VAE_DECODER_OV_PATH_INPAINT, "CPU") - vae_encoder_inpaint = core.compile_model(VAE_ENCODER_OV_PATH_INPAINT, "CPU") + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=2, options=('CPU', 'GPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + + text_enc_inpaint = core.compile_model(TEXT_ENCODER_OV_PATH_INPAINT, device.value) + unet_model_inpaint = core.compile_model(UNET_OV_PATH_INPAINT, device.value) + vae_decoder_inpaint = core.compile_model(VAE_DECODER_OV_PATH_INPAINT, device.value) + vae_encoder_inpaint = core.compile_model(VAE_ENCODER_OV_PATH_INPAINT, device.value) ov_pipe_inpaint = OVStableDiffusionInpaintingPipeline( tokenizer=tokenizer, @@ -1165,8 +1322,9 @@ Configuration steps: scheduler=scheduler_inpaint, ) -Run Infinite Zoom video generation -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Run Infinite Zoom video generation `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -1215,3 +1373,18 @@ Run Infinite Zoom video generation ) ipaddr = gethostbyname(gethostname()) demo.queue().launch(share=True) + + +.. parsed-literal:: + + Running on local URL: http://127.0.0.1:7861 + Running on public URL: https://462b1833bf3b980731.gradio.live + + This share link expires in 72 hours. For free permanent hosting and GPU upgrades, run `gradio deploy` from Terminal to deploy to Spaces (https://huggingface.co/spaces) + + + +.. raw:: html + +
+ diff --git a/docs/notebooks/236-stable-diffusion-v2-optimum-demo-comparison-with-output.rst b/docs/notebooks/236-stable-diffusion-v2-optimum-demo-comparison-with-output.rst index 094c2574882..015ce99aac8 100644 --- a/docs/notebooks/236-stable-diffusion-v2-optimum-demo-comparison-with-output.rst +++ b/docs/notebooks/236-stable-diffusion-v2-optimum-demo-comparison-with-output.rst @@ -1,6 +1,18 @@ Stable Diffusion v2.1 using Optimum-Intel OpenVINO and multiple Intel Hardware ============================================================================== +.. _top: + +|image0| + +**Table of contents**: + +- `Showing Info Available Devices <#showing-info-available-devices>`__ +- `Using full precision model in CPU with StableDiffusionPipeline <#using-full-precision-model-in-cpu-with-stablediffusionpipeline>`__ +- `Using full precision model in CPU with OVStableDiffusionPipeline <#using-full-precision-model-in-cpu-with-ovstablediffusionpipeline>`__ +- `Using full precision model in dGPU with OVStableDiffusionPipeline <#using-full-precision-model-in-dgpu-with-ovstablediffusionpipeline>`__ + +.. |image0| image:: https://github.com/openvinotoolkit/openvino_notebooks/assets/10940214/1858dae4-72fd-401e-b055-66d503d82446 Optimum Intel is the interface between the Transformers and Diffusers libraries and the different tools and libraries provided by Intel to @@ -19,8 +31,9 @@ this import warnings warnings.filterwarnings('ignore') -Showing Info Available Devices -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Showing Info Available Devices `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The ``available_devices`` property shows the available devices in your system. The “FULL_DEVICE_NAME” option to ``ie.get_property()`` shows the @@ -31,7 +44,7 @@ If you just have either an iGPU or dGPU that will be assigned to ``"GPU"`` Note: For more details about GPU with OpenVINO visit this -`link `__. +`link `__. If you have been facing any issue in Ubuntu 20.04 or Windows 11 read this `blog `__. @@ -54,8 +67,8 @@ this GPU: Intel(R) Data Center GPU Flex 170 (dGPU) -Using full precision model in CPU with StableDiffusionPipeline -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Using full precision model in CPU with ``StableDiffusionPipeline``. `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ .. code:: ipython3 @@ -106,8 +119,9 @@ Using full precision model in CPU with StableDiffusionPipeline -Using full precision model in CPU with OVStableDiffusionPipeline -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Using full precision model in CPU with ``OVStableDiffusionPipeline``. `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -195,8 +209,9 @@ Using full precision model in CPU with OVStableDiffusionPipeline -Using full precision model in dGPU with OVStableDiffusionPipeline -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Using full precision model in dGPU with ``OVStableDiffusionPipeline``. `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The model in this notebook is FP32 precision. And thanks to the new feature of OpenVINO 2023.0 you do not need to convert the model to FP16 diff --git a/docs/notebooks/236-stable-diffusion-v2-optimum-demo-comparison-with-output_files/index.html b/docs/notebooks/236-stable-diffusion-v2-optimum-demo-comparison-with-output_files/index.html index 78ec5a94b12..f88c1f722ef 100644 --- a/docs/notebooks/236-stable-diffusion-v2-optimum-demo-comparison-with-output_files/index.html +++ b/docs/notebooks/236-stable-diffusion-v2-optimum-demo-comparison-with-output_files/index.html @@ -1,9 +1,9 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/236-stable-diffusion-v2-optimum-demo-comparison-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/236-stable-diffusion-v2-optimum-demo-comparison-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/236-stable-diffusion-v2-optimum-demo-comparison-with-output_files/


../
-236-stable-diffusion-v2-optimum-demo-comparison..> 12-Jul-2023 00:11              573225
-236-stable-diffusion-v2-optimum-demo-comparison..> 12-Jul-2023 00:11              569855
-236-stable-diffusion-v2-optimum-demo-comparison..> 12-Jul-2023 00:11              466925
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/236-stable-diffusion-v2-optimum-demo-comparison-with-output_files/


../
+236-stable-diffusion-v2-optimum-demo-comparison..> 16-Aug-2023 01:31              573225
+236-stable-diffusion-v2-optimum-demo-comparison..> 16-Aug-2023 01:31              569855
+236-stable-diffusion-v2-optimum-demo-comparison..> 16-Aug-2023 01:31              466925
 

diff --git a/docs/notebooks/236-stable-diffusion-v2-optimum-demo-with-output.rst b/docs/notebooks/236-stable-diffusion-v2-optimum-demo-with-output.rst index 8b33b5910d6..b3e4d3e49dd 100644 --- a/docs/notebooks/236-stable-diffusion-v2-optimum-demo-with-output.rst +++ b/docs/notebooks/236-stable-diffusion-v2-optimum-demo-with-output.rst @@ -1,6 +1,19 @@ Stable Diffusion v2.1 using Optimum-Intel OpenVINO ================================================== +.. _top: + +|image0| + +**Table of contents**: + +- `Showing Info Available Devices <#showing-info-available-devices>`__ +- `Download Pre-Converted Stable Diffusion 2.1 IR <#download-pre-converted-stable-diffusion-2.1-ir>`__ +- `Save the pre-trained models, Select the inference device and compile it <#save-the-pre-trained-models-select-the-inference-device-and-compile-it>`__ +- `Be creative, add the prompt and enjoy the result <#be-creative-add-the-prompt-and-enjoy-the-result>`__ + +.. |image0| image:: https://github.com/openvinotoolkit/openvino_notebooks/assets/10940214/1858dae4-72fd-401e-b055-66d503d82446 + Optimum Intel is the interface between the Transformers and Diffusers libraries and the different tools and libraries provided by Intel to accelerate end-to-end pipelines on Intel architectures. More details in @@ -29,19 +42,20 @@ Autoencoder with Decoder and Encoder models. image The base model used for this example is the -“`stabilityai/stable-diffusion-2-1-base `__. +`stabilityai/stable-diffusion-2-1-base `__. This model was converted to OpenVINO format, for accelerated inference on CPU or Intel GPU with OpenVINO’s integration into Optimum: -optimum-intel. The model weights are stored with FP16 precision, which -reduces the size of the model by half. You can find the model used in -this notebook -is”\ `helenai/stabilityai-stable-diffusion-2-1-base-ov `__". +``optimum-intel``. The model weights are stored with FP16 precision, +which reduces the size of the model by half. You can find the model used +in this notebook is +`helenai/stabilityai-stable-diffusion-2-1-base-ov `__. Let’s download the pre-converted model Stable Diffusion 2.1 `Intermediate Representation Format (IR) `__ -Showing Info Available Devices -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Showing Info Available Devices `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The ``available_devices`` property shows the available devices in your system. The “FULL_DEVICE_NAME” option to ``ie.get_property()`` shows the @@ -52,7 +66,7 @@ If you just have either an iGPU or dGPU that will be assigned to ``"GPU"`` Note: For more details about GPU with OpenVINO visit this -`link `__. +`link `__. If you have been facing any issue in Ubuntu 20.04 or Windows 11 read this `blog `__. @@ -71,12 +85,14 @@ this .. parsed-literal:: - CPU: Intel(R) Core(TM) i9-10980XE CPU @ 3.00GHz - GPU: NVIDIA GeForce GTX 1080 Ti (dGPU) + CPU: 13th Gen Intel(R) Core(TM) i9-13900K + GPU.0: Intel(R) UHD Graphics 770 (iGPU) + GPU.1: Intel(R) Arc(TM) A770 Graphics (dGPU) -Download Pre-Converted Stable Diffusion 2.1 IR -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Download Pre-Converted Stable Diffusion 2.1 IR `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -166,8 +182,9 @@ Download Pre-Converted Stable Diffusion 2.1 IR -Save the pre-trained models, Select the inference device and compile it -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Save the pre-trained models, Select the inference device and compile it. `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + You can save the model locally in order to avoid downloading process later. The model will also saved in the cache. @@ -186,8 +203,9 @@ later. The model will also saved in the cache. Compiling the unet... -Be creative, add the prompt and enjoy the result -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Be creative, add the prompt and enjoy the result `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 diff --git a/docs/notebooks/236-stable-diffusion-v2-optimum-demo-with-output_files/index.html b/docs/notebooks/236-stable-diffusion-v2-optimum-demo-with-output_files/index.html index 9aff6f440b2..34231ab3f03 100644 --- a/docs/notebooks/236-stable-diffusion-v2-optimum-demo-with-output_files/index.html +++ b/docs/notebooks/236-stable-diffusion-v2-optimum-demo-with-output_files/index.html @@ -1,7 +1,7 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/236-stable-diffusion-v2-optimum-demo-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/236-stable-diffusion-v2-optimum-demo-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/236-stable-diffusion-v2-optimum-demo-with-output_files/


../
-236-stable-diffusion-v2-optimum-demo-with-outpu..> 12-Jul-2023 00:11              451596
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/236-stable-diffusion-v2-optimum-demo-with-output_files/


../
+236-stable-diffusion-v2-optimum-demo-with-outpu..> 16-Aug-2023 01:31              451596
 

diff --git a/docs/notebooks/236-stable-diffusion-v2-text-to-image-demo-with-output.rst b/docs/notebooks/236-stable-diffusion-v2-text-to-image-demo-with-output.rst index 36f0a97fa87..2c54a64fc99 100644 --- a/docs/notebooks/236-stable-diffusion-v2-text-to-image-demo-with-output.rst +++ b/docs/notebooks/236-stable-diffusion-v2-text-to-image-demo-with-output.rst @@ -1,6 +1,8 @@ Stable Diffusion Text-to-Image Demo =================================== +.. _top: + Stable Diffusion is an innovative generative AI technique that allows us to generate and manipulate images in interesting ways, including generating image from text and restoring missing parts of pictures @@ -12,17 +14,32 @@ less restrictive filtering of the dataset. All of these features give us promising results for selecting a wide range of input text prompts! **Note:** This is a shorter version of the -`236-stable-diffusion-v2-text-to-image.ipynb `__ +`236-stable-diffusion-v2-text-to-image `__ notebook for demo purposes and to get started quickly. This version does not have the full implementation of the helper utilities needed to convert the models from PyTorch to ONNX to OpenVINO, and the OpenVINO -OVStableDiffusion pipeline within the notebook directly. If you would +``OVStableDiffusionPipeline`` within the notebook directly. If you would like to see the full implementation of stable diffusion for text to image, please visit -`236-stable-diffusion-v2-text-to-image.ipynb `__. +`236-stable-diffusion-v2-text-to-image `__. + +**Table of contents**: + +- `Step 0: Install and import prerequisites <#step-0-install-and-import-prerequisites>`__ +- `Step 1: Stable Diffusion v2 Fundamental components <#step-1-stable-diffusion-v2-fundamental-components>`__ + + - `Step 1.1: Retrieve components from HuggingFace <#step-1-1-retrieve-components-from-huggingface>`__ + +- `Step 2: Convert the models to OpenVINO <#step-2-convert-the-models-to-openvino>`__ +- `Step 3: Text-to-Image Generation Inference Pipeline <#step-3-text-to-image-generation-inference-pipeline>`__ + + - `Step 3.1: Load and Understand Text to Image OpenVINO models <#step-3-1-load-and-understand-text-to-image-openvino-models>`__ + - `Select inference device <#select-inference-device>`__ + - `Step 3.3: Run Text-to-Image generation <#step-3-3-run-text-to-image-generation>`__ + +Step 0: Install and import prerequisites `⇑ <#top>`__ +############################################################################################################################### -Step 0: Install and import prerequisites ----------------------------------------- .. code:: ipython3 @@ -33,15 +50,16 @@ To work with Stable Diffusion v2, we will use Hugging Face’s To experiment with Stable Diffusion models, Diffusers exposes the `StableDiffusionPipeline `__ -and StableDiffusionInpaintPipeline, similar to the `other Diffusers +and ``StableDiffusionInpaintPipeline``, similar to the `other Diffusers pipelines `__. .. code:: ipython3 !pip install -q "diffusers>=0.14.0" openvino-dev openvino "transformers >= 4.25.1" accelerate -Step 1: Stable Diffusion v2 Fundamental components --------------------------------------------------- +Step 1: Stable Diffusion v2 Fundamental components `⇑ <#top>`__ +############################################################################################################################### + Stable Diffusion pipelines for both Text to Image and Inpainting consist of three important parts: @@ -52,11 +70,12 @@ of three important parts: 2. A U-Net for step-by-step denoising of latent image representation. 3. An Autoencoder (VAE) for decoding the latent space to an image. -Depending on the pipeline, the paramaters for these parts can differ, +Depending on the pipeline, the parameters for these parts can differ, which we’ll explore in this demo! -Step 1.1: Retrieve components from HuggingFace -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Step 1.1: Retrieve components from HuggingFace `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Let’s start by retrieving these components from HuggingFace! @@ -88,8 +107,9 @@ using ``stable-diffusion-2-1``. text_encoder\model.safetensors not found -Step 2: Convert the models to OpenVINO --------------------------------------- +Step 2: Convert the models to OpenVINO `⇑ <#top>`__ +############################################################################################################################### + Now that we’ve retrieved the three parts for both of these pipelines, we now need to: @@ -102,8 +122,8 @@ now need to: torch.onnx.export(model_part, image, onnx_path, input_names=[ '...'], output_names=['...']) -2. Convert these ONNX models to OpenVINO IR format, using Model - Optimizer tool using: +2. Convert these ONNX models to OpenVINO IR format, using ``mo`` + command-line tool: :: @@ -147,29 +167,49 @@ pipelines in OpenVINO on our own data! WARNING:root:Failed to send event with error cannot schedule new futures after shutdown. -3. Text-to-Image Generation Inference Pipeline ----------------------------------------------- +Step 3: Text-to-Image Generation Inference Pipeline `⇑ <#top>`__ +############################################################################################################################### -Step 3.1: Load and Understand Text to Image OpenVINO models -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ -First, let’s create instances of our OpenVINO Model for Text to Image. +Step 3.1: Load and Understand Text to Image OpenVINO models `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: .. code:: ipython3 + import ipywidgets as widgets from openvino.runtime import Core core = Core() - text_enc = core.compile_model(txt_encoder_ov_path, "CPU") + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + +Let’s create instances of our OpenVINO Model for Text to Image. .. code:: ipython3 - unet_model = core.compile_model(unet_ov_path, 'CPU') + text_enc = core.compile_model(txt_encoder_ov_path, device.value) .. code:: ipython3 - vae_encoder = core.compile_model(vae_encoder_ov_path, 'CPU') - vae_decoder = core.compile_model(vae_decoder_ov_path, 'CPU') + unet_model = core.compile_model(unet_ov_path, device.value) + +.. code:: ipython3 + + vae_encoder = core.compile_model(vae_encoder_ov_path, device.value) + vae_decoder = core.compile_model(vae_decoder_ov_path, device.value) Next, we will define a few key elements to create the inference pipeline, as depicted in the diagram below: @@ -179,7 +219,7 @@ pipeline, as depicted in the diagram below: text2img-stable-diffusion -As part of the OVStableDiffusionPipeline() class: +As part of the ``OVStableDiffusionPipeline()`` class: 1. The stable diffusion pipeline takes both a latent seed and a text prompt as input. The latent seed is used to generate random latent @@ -190,7 +230,7 @@ As part of the OVStableDiffusionPipeline() class: representations while being conditioned on the text embeddings. The output of the U-Net, being the noise residual, is used to compute a denoised latent image representation via a scheduler algorithm. In - this case we use the LMSDiscrete scheduler. + this case we use the ``LMSDiscreteScheduler``. .. code:: ipython3 @@ -217,8 +257,9 @@ As part of the OVStableDiffusionPipeline() class: from diffusers.pipeline_utils import DiffusionPipeline -Step 3.3: Run Text-to-Image generation -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Step 3.3: Run Text-to-Image generation `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Now, let’s define some text prompts for image generation and run our inference pipeline. @@ -237,7 +278,7 @@ To improve image generation quality, we can use negative prompting. While positive prompts steer diffusion toward the images associated with it, negative prompts declares undesired concepts for the generation image, e.g. if we want to have colorful and bright images, a gray scale -image will be result which we want to avoid. In this case, a grey scale +image will be result which we want to avoid. In this case, a gray scale can be treated as negative prompt. The positive and negative prompt are in equal footing. You can always use one with or without the other. More explanation of how it works can be found in this @@ -284,6 +325,6 @@ explanation of how it works can be found in this -.. image:: 236-stable-diffusion-v2-text-to-image-demo-with-output_files/236-stable-diffusion-v2-text-to-image-demo-with-output_23_0.png +.. image:: 236-stable-diffusion-v2-text-to-image-demo-with-output_files/236-stable-diffusion-v2-text-to-image-demo-with-output_25_0.png diff --git a/docs/notebooks/236-stable-diffusion-v2-text-to-image-demo-with-output_files/236-stable-diffusion-v2-text-to-image-demo-with-output_23_0.png b/docs/notebooks/236-stable-diffusion-v2-text-to-image-demo-with-output_files/236-stable-diffusion-v2-text-to-image-demo-with-output_25_0.png similarity index 100% rename from docs/notebooks/236-stable-diffusion-v2-text-to-image-demo-with-output_files/236-stable-diffusion-v2-text-to-image-demo-with-output_23_0.png rename to docs/notebooks/236-stable-diffusion-v2-text-to-image-demo-with-output_files/236-stable-diffusion-v2-text-to-image-demo-with-output_25_0.png diff --git a/docs/notebooks/236-stable-diffusion-v2-text-to-image-demo-with-output_files/index.html b/docs/notebooks/236-stable-diffusion-v2-text-to-image-demo-with-output_files/index.html index a7f747bc4c6..63ddd649d42 100644 --- a/docs/notebooks/236-stable-diffusion-v2-text-to-image-demo-with-output_files/index.html +++ b/docs/notebooks/236-stable-diffusion-v2-text-to-image-demo-with-output_files/index.html @@ -1,7 +1,7 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/236-stable-diffusion-v2-text-to-image-demo-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/236-stable-diffusion-v2-text-to-image-demo-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/236-stable-diffusion-v2-text-to-image-demo-with-output_files/


../
-236-stable-diffusion-v2-text-to-image-demo-with..> 12-Jul-2023 00:11             1009164
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/236-stable-diffusion-v2-text-to-image-demo-with-output_files/


../
+236-stable-diffusion-v2-text-to-image-demo-with..> 16-Aug-2023 01:31             1009164
 

diff --git a/docs/notebooks/236-stable-diffusion-v2-text-to-image-with-output.rst b/docs/notebooks/236-stable-diffusion-v2-text-to-image-with-output.rst index ebf7747262c..e1196e0625a 100644 --- a/docs/notebooks/236-stable-diffusion-v2-text-to-image-with-output.rst +++ b/docs/notebooks/236-stable-diffusion-v2-text-to-image-with-output.rst @@ -1,6 +1,8 @@ Text-to-Image Generation with Stable Diffusion v2 and OpenVINO™ =============================================================== +.. _top: + Stable Diffusion v2 is the next generation of Stable Diffusion model a Text-to-Image latent diffusion model created by the researchers and engineers from `Stability AI `__ and @@ -33,12 +35,33 @@ by the other models that have emerged since the introduction of the first iteration. Some of the features that can be found in the new model are: -* The model comes with a new robust encoder, OpenCLIP, created by LAION and aided by Stability AI; this version v2 significantly enhances the produced photos over the V1 versions. -* The model can now generate images in a 768x768 resolution, offering more information to be shown in the generated images. -* The model finetuned with `v-objective `__. The v-parameterization is particularly useful for numerical stability throughout the diffusion process to enable progressive distillation for models. For models that operate at higher resolution, it is also discovered that the v-parameterization avoids color shifting artifacts that are known to affect high resolution diffusion models, and in the video setting it avoids temporal color shifting that sometimes appears with epsilon-prediction used in Stable Diffusion v1. -* The model also comes with a new diffusion model capable of running upscaling on the images generated. Upscaled images can be adjusted up to 4 times the original image. Provided as separated model, for more details please check `stable-diffusion-x4-upscaler `__ -* The model comes with a new refined depth architecture capable of preserving context from prior generation layers in an img2img setting. This structure preservation helps generate images that preserving forms and shadow of objects, but with different content. -* The model comes with an updated inpainting module built upon the previous model. This text-guided inpainting makes switching out parts in the image easier than before. +- The model comes with a new robust encoder, OpenCLIP, created by LAION + and aided by Stability AI; this version v2 significantly enhances the + produced photos over the V1 versions. +- The model can now generate images in a 768x768 resolution, offering + more information to be shown in the generated images. +- The model finetuned with + `v-objective `__. The + v-parameterization is particularly useful for numerical stability + throughout the diffusion process to enable progressive distillation + for models. For models that operate at higher resolution, it is also + discovered that the v-parameterization avoids color shifting + artifacts that are known to affect high resolution diffusion models, + and in the video setting it avoids temporal color shifting that + sometimes appears with epsilon-prediction used in Stable Diffusion + v1. +- The model also comes with a new diffusion model capable of running + upscaling on the images generated. Upscaled images can be adjusted up + to 4 times the original image. Provided as separated model, for more + details please check + `stable-diffusion-x4-upscaler `__ +- The model comes with a new refined depth architecture capable of + preserving context from prior generation layers in an image-to-image + setting. This structure preservation helps generate images that + preserving forms and shadow of objects, but with different content. +- The model comes with an updated inpainting module built upon the + previous model. This text-guided inpainting makes switching out parts + in the image easier than before. This notebook demonstrates how to convert and run Stable Diffusion v2 model using OpenVINO. @@ -46,18 +69,33 @@ model using OpenVINO. Notebook contains the following steps: 1. Convert PyTorch models to ONNX format. -2. Convert ONNX models to OpenVINO IR format, using Model Optimizer tool. +2. Convert ONNX models to OpenVINO IR format, using model conversion + API. 3. Run Stable Diffusion v2 Text-to-Image pipeline with OpenVINO. **Note:** This is the full version of the Stable Diffusion text-to-image implementation. If you would like to get started and run the notebook -quickly, check out -`236-stable-diffusion-v2-text-to-image-demo.ipynb `__. +quickly, check out `236-stable-diffusion-v2-text-to-image-demo +notebook `__. -Prerequisites -------------- +**Table of contents**: -install required packages +- `Prerequisites <#prerequisites>`__ +- `Stable Diffusion v2 for Text-to-Image Generation <#stable-diffusion-v2-for-text-to-image-generation>`__ + + - `Stable Diffusion in Diffusers library <#stable-diffusion-in-diffusers-library>`__ + - `Convert models to OpenVINO Intermediate representation (IR) format <#convert-models-to-openvino-intermediate-representation-ir-format>`__ + - `Text Encoder <#text-encoder>`__ + - `U-Net <#u-net>`__ + - `VAE <#vae>`__ + - `Prepare Inference Pipeline <#prepare-inference-pipeline>`__ + - `Configure Inference Pipeline <#configure-inference-pipeline>`__ + - `Run Text-to-Image generation <#run-text-to-image-generation>`__ + +Prerequisites `⇑ <#top>`__ +############################################################################################################################### + +Install required packages: .. code:: ipython3 @@ -67,18 +105,19 @@ install required packages .. parsed-literal:: - [notice] A new release of pip available: 22.3.1 -> 23.0.1 + [notice] A new release of pip is available: 23.1.2 -> 23.2 [notice] To update, run: pip install --upgrade pip -Stable Diffusion v2 for Text-to-Image Generation ------------------------------------------------- +Stable Diffusion v2 for Text-to-Image Generation `⇑ <#top>`__ +############################################################################################################################### + To start, let’s look on Text-to-Image process for Stable Diffusion v2. -We will use -`stabilitai/stable-diffusion-2-1 `__ -model for these purposes. The main difference from Stable Diffusion v2 -and Stable Diffusion v2.1 is usage of more data, more training, and less +We will use `Stable Diffusion +v2-1 `__ model +for these purposes. The main difference from Stable Diffusion v2 and +Stable Diffusion v2.1 is usage of more data, more training, and less restrictive filtering of the dataset, that gives promising results for selecting wide range of input text prompts. More details about model can be found in `Stability AI blog @@ -86,8 +125,8 @@ post `__ and original model `repository `__. -Stable Diffusion in Diffusers library -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Stable Diffusion in Diffusers library `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ To work with Stable Diffusion v2, we will use Hugging Face `Diffusers `__ library. To @@ -117,20 +156,17 @@ using ``stable-diffusion-2-1``: del pipe - .. parsed-literal:: - Fetching 13 files: 0%| | 0/13 [00:00`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ - /home/ea/work/transformers/src/transformers/models/clip/feature_extraction_clip.py:28: FutureWarning: The class CLIPFeatureExtractor is deprecated and will be removed in version 5 of Transformers. Please use CLIPImageProcessor instead. - warnings.warn( - - -Convert models to OpenVINO Intermediate representation (IR) format -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ OpenVINO supports PyTorch through export to the ONNX format. We will use the ``torch.onnx.export`` function to obtain the ONNX model, we can @@ -149,14 +185,16 @@ to convert a model to IR format. The pipeline consists of three important parts: -* Text Encoder to create condition to generate an image from a text prompt. -* U-Net for step-by-step denoising latent image representation. -* Autoencoder (VAE) for decoding latent space to image. +- Text Encoder to create condition to generate an image from a text + prompt. +- U-Net for step-by-step denoising latent image representation. +- Autoencoder (VAE) for decoding latent space to image. Let us convert each part: -Text Encoder -~~~~~~~~~~~~ +Text Encoder `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The text-encoder is responsible for transforming the input prompt, for example, “a photo of an astronaut riding a horse” into an embedding @@ -232,45 +270,22 @@ this opset. .. parsed-literal:: - /tmp/ipykernel_383583/1233802758.py:26: FutureWarning: 'torch.onnx._export' is deprecated in version 1.12.0 and will be removed in version 1.14. Please use `torch.onnx.export` instead. - torch.onnx._export( - /home/ea/work/transformers/src/transformers/models/clip/modeling_clip.py:759: TracerWarning: torch.tensor results are registered as constants in the trace. You can safely ignore this warning if you use this function to create tensors out of constant variables that would be the same every time you call this function. In any other case, this might cause the trace to be incorrect. - mask.fill_(torch.tensor(torch.finfo(dtype).min)) - /home/ea/work/transformers/src/transformers/models/clip/modeling_clip.py:284: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! - if attn_weights.size() != (bsz * self.num_heads, tgt_len, src_len): - /home/ea/work/transformers/src/transformers/models/clip/modeling_clip.py:292: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! - if causal_attention_mask.size() != (bsz, 1, tgt_len, src_len): - /home/ea/work/transformers/src/transformers/models/clip/modeling_clip.py:324: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! - if attn_output.size() != (bsz * self.num_heads, tgt_len, self.head_dim): - /home/ea/work/notebooks_env/lib/python3.8/site-packages/torch/onnx/symbolic_helper.py:710: UserWarning: Type cannot be inferred, which might cause exported graph to produce incorrect results. - warnings.warn( - /home/ea/work/notebooks_env/lib/python3.8/site-packages/torch/onnx/symbolic_opset9.py:5408: UserWarning: Exporting aten::index operator of advanced indexing in opset 14 is achieved by combination of multiple ONNX operators, including Reshape, Transpose, Concat, and Gather. If indices include negative values, the exported graph will produce incorrect results. - warnings.warn( + Text encoder will be loaded from sd2.1/text_encoder.xml -.. parsed-literal:: +U-Net `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ - Text Encoder successfully converted to ONNX - Warning: One or more of the values of the Constant can't fit in the float16 data type. Those values were casted to the nearest limit value, the model can produce incorrect results. - [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. - Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html - [ SUCCESS ] Generated IR version 11 model. - [ SUCCESS ] XML file: /home/ea/work/openvino_notebooks/notebooks/236-stable-diffusion-v2/sd2.1/text_encoder.xml - [ SUCCESS ] BIN file: /home/ea/work/openvino_notebooks/notebooks/236-stable-diffusion-v2/sd2.1/text_encoder.bin - Text Encoder successfully converted to IR - - -U-Net -~~~~~ U-Net model gradually denoises latent image representation guided by text encoder hidden state. -U-Net model has three inputs: +U-Net model has three inputs: -* ``sample`` - latent image sample from previous step. Generation process has not been started yet, so you will use random noise. -* ``timestep`` - current scheduler step. -* ``encoder_hidden_state`` - hidden state of text encoder. +- ``sample`` - latent image sample from previous step. Generation + process has not been started yet, so you will use random noise. +- ``timestep`` - current scheduler step. +- ``encoder_hidden_state`` - hidden state of text encoder. Model predicts the ``sample`` state for the next step. @@ -280,7 +295,7 @@ pretrained to generate images with resolution 768x768, initial latent sample size for this case is 96x96. Besides that, for different use cases like inpainting and depth to image generation model also can accept additional image information: depth map or mask as channel-wise -concatenation with initial latent sample. For convering U-Net model for +concatenation with initial latent sample. For converting U-Net model for such use cases required to modify number of input channels. .. code:: ipython3 @@ -340,35 +355,12 @@ such use cases required to modify number of input channels. .. parsed-literal:: - /tmp/ipykernel_383583/4211352295.py:32: FutureWarning: 'torch.onnx._export' is deprecated in version 1.12.0 and will be removed in version 1.14. Please use `torch.onnx.export` instead. - torch.onnx._export( - /home/ea/work/diffusers/src/diffusers/models/unet_2d_condition.py:526: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! - if any(s % default_overall_up_factor != 0 for s in sample.shape[-2:]): - /home/ea/work/diffusers/src/diffusers/models/resnet.py:185: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! - assert hidden_states.shape[1] == self.channels - /home/ea/work/diffusers/src/diffusers/models/resnet.py:190: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! - assert hidden_states.shape[1] == self.channels - /home/ea/work/diffusers/src/diffusers/models/resnet.py:112: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! - assert hidden_states.shape[1] == self.channels - /home/ea/work/diffusers/src/diffusers/models/resnet.py:125: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! - if hidden_states.shape[0] >= 64: - /home/ea/work/diffusers/src/diffusers/models/unet_2d_condition.py:651: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! - if not return_dict: + U-Net will be loaded from sd2.1/unet.xml -.. parsed-literal:: +VAE `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ - U-Net successfully converted to ONNX - [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. - Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html - [ SUCCESS ] Generated IR version 11 model. - [ SUCCESS ] XML file: /home/ea/work/openvino_notebooks/notebooks/236-stable-diffusion-v2/sd2.1/unet.xml - [ SUCCESS ] BIN file: /home/ea/work/openvino_notebooks/notebooks/236-stable-diffusion-v2/sd2.1/unet.bin - U-Net successfully converted to IR - - -VAE -~~~ The VAE model has two parts, an encoder and a decoder. The encoder is used to convert the image into a low dimensional latent representation, @@ -487,48 +479,13 @@ RAM (recommended at least 32GB). .. parsed-literal:: - /home/ea/work/notebooks_env/lib/python3.8/site-packages/torch/onnx/_internal/jit_utils.py:258: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) - _C._jit_pass_onnx_node_shape_type_inference(node, params_dict, opset_version) - /home/ea/work/notebooks_env/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) - _C._jit_pass_onnx_graph_shape_type_inference( - /home/ea/work/notebooks_env/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) - _C._jit_pass_onnx_graph_shape_type_inference( + VAE encoder will be loaded from sd2.1/vae_encoder.xml + VAE decoder will be loaded from sd2.1/vae_decoder.xml -.. parsed-literal:: +Prepare Inference Pipeline `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ - VAE encoder successfully converted to ONNX - [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. - Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html - [ SUCCESS ] Generated IR version 11 model. - [ SUCCESS ] XML file: /home/ea/work/openvino_notebooks/notebooks/236-stable-diffusion-v2/sd2.1/vae_encoder.xml - [ SUCCESS ] BIN file: /home/ea/work/openvino_notebooks/notebooks/236-stable-diffusion-v2/sd2.1/vae_encoder.bin - VAE encoder successfully converted to IR - - -.. parsed-literal:: - - /home/ea/work/notebooks_env/lib/python3.8/site-packages/torch/onnx/_internal/jit_utils.py:258: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) - _C._jit_pass_onnx_node_shape_type_inference(node, params_dict, opset_version) - /home/ea/work/notebooks_env/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) - _C._jit_pass_onnx_graph_shape_type_inference( - /home/ea/work/notebooks_env/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) - _C._jit_pass_onnx_graph_shape_type_inference( - - -.. parsed-literal:: - - VAE decoder successfully converted to ONNX - [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. - Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html - [ SUCCESS ] Generated IR version 11 model. - [ SUCCESS ] XML file: /home/ea/work/openvino_notebooks/notebooks/236-stable-diffusion-v2/sd2.1/vae_decoder.xml - [ SUCCESS ] BIN file: /home/ea/work/openvino_notebooks/notebooks/236-stable-diffusion-v2/sd2.1/vae_decoder.bin - VAE decoder successfully converted to IR - - -Prepare Inference Pipeline -~~~~~~~~~~~~~~~~~~~~~~~~~~ Putting it all together, let us now take a closer look at how the model works in inference by illustrating the logical flow. @@ -570,9 +527,20 @@ The chart above looks very similar to Stable Diffusion V1 from `notebook <225-stable-diffusion-text-to-image-with-output.html>`__, but there is some small difference in details: -* Changed input resolution for U-Net model. -* Changed text encoder and as the result size of its hidden state embeddings. -* Additionally, to improve image generation quality authors introduced negative prompting. Technically, positive prompt steers the diffusion toward the images associated with it, while negative prompt steers the diffusion away from it.In other words, negative prompt declares undesired concepts for generation image, e.g. if we want to have colorful and bright image, gray scale image will be result which we want to avoid, in this case grey scale can be treated as negative prompt. The positive and negative prompt are in equal footing. You can always use one with or without the other. More explanation of how it works can be found in this `article `__. +- Changed input resolution for U-Net model. +- Changed text encoder and as the result size of its hidden state + embeddings. +- Additionally, to improve image generation quality authors introduced + negative prompting. Technically, positive prompt steers the diffusion + toward the images associated with it, while negative prompt steers + the diffusion away from it.In other words, negative prompt declares + undesired concepts for generation image, e.g. if we want to have + colorful and bright image, gray scale image will be result which we + want to avoid, in this case gray scale can be treated as negative + prompt. The positive and negative prompt are in equal footing. You + can always use one with or without the other. More explanation of how + it works can be found in this + `article `__. .. code:: ipython3 @@ -924,19 +892,50 @@ but there is some small difference in details: return timesteps, num_inference_steps - t_start -Configure Inference Pipeline -~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +.. parsed-literal:: + + /tmp/ipykernel_1185037/1028096992.py:9: FutureWarning: Importing `DiffusionPipeline` or `ImagePipelineOutput` from diffusers.pipeline_utils is deprecated. Please import from diffusers.pipelines.pipeline_utils instead. + from diffusers.pipeline_utils import DiffusionPipeline + + +Configure Inference Pipeline `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + First, you should create instances of OpenVINO Model. .. code:: ipython3 + import ipywidgets as widgets from openvino.runtime import Core + core = Core() - text_enc = core.compile_model(TEXT_ENCODER_OV_PATH, "CPU") - unet_model = core.compile_model(UNET_OV_PATH, 'CPU') - vae_decoder = core.compile_model(VAE_DECODER_OV_PATH, 'CPU') - vae_encoder = core.compile_model(VAE_ENCODER_OV_PATH, 'CPU') + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=2, options=('CPU', 'GPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + core = Core() + text_enc = core.compile_model(TEXT_ENCODER_OV_PATH, device.value) + unet_model = core.compile_model(UNET_OV_PATH, device.value) + vae_decoder = core.compile_model(VAE_DECODER_OV_PATH, device.value) + vae_encoder = core.compile_model(VAE_ENCODER_OV_PATH, device.value) Model tokenizer and scheduler are also important parts of the pipeline. Let us define them and put all components together. @@ -957,21 +956,20 @@ Let us define them and put all components together. scheduler=scheduler ) -Run Text-to-Image generation -~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Run Text-to-Image generation `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Now, you can define a text prompts for image generation and run inference pipeline. Optionally, you can also change the random generator -seed for latent state initialization and number of steps. +seed for latent state initialization and number of steps. -.. note:: - - Consider increasing ``steps`` to get more precise results. A suggested value is ``50``, but it will take longer time to process. + **Note**: Consider increasing ``steps`` to get more precise results. + A suggested value is ``50``, but it will take longer time to process. .. code:: ipython3 import gradio as gr - from socket import gethostbyname, gethostname def generate(prompt, negative_prompt, seed, num_steps, _=gr.Progress(track_tqdm=True)): @@ -1001,5 +999,22 @@ seed for latent state initialization and number of steps. ], "image", ) - ipaddr = gethostbyname(gethostname()) - demo.queue().launch(server_name=ipaddr) + + try: + demo.queue().launch() + except Exception: + demo.queue().launch(share=True) + + +.. parsed-literal:: + + Running on local URL: http://127.0.0.1:7863 + + To create a public link, set `share=True` in `launch()`. + + + +.. raw:: html + +
+ diff --git a/docs/notebooks/237-segment-anything-with-output.rst b/docs/notebooks/237-segment-anything-with-output.rst index c344e76e969..a8109bf23ee 100644 --- a/docs/notebooks/237-segment-anything-with-output.rst +++ b/docs/notebooks/237-segment-anything-with-output.rst @@ -1,6 +1,36 @@ Object masks from prompts with SAM and OpenVINO =============================================== +.. _top: + +**Table of contents**: + +- `Background <#background>`__ +- `Prerequisites <#prerequisites>`__ +- `Convert model to OpenVINO Intermediate Representation <#convert-model-to-openvino-intermediate-representation>`__ + + - `Download model checkpoint and create PyTorch model <#download-model-checkpoint-and-create-pytorch-model>`__ + - `Image Encoder <#image-encoder>`__ + - `Mask predictor <#mask-predictor>`__ + +- `Run OpenVINO model in interactive segmentation mode <#run-openvino-model-in-interactive-segmentation-mode>`__ + + - `Example Image <#example-image>`__ + - `Preprocessing and visualization utilities <#preprocessing-and-visualization-utilities>`__ + - `Image encoding <#image-encoding>`__ + - `Example point input <#example-point-input>`__ + - `Example with multiple points <#example-with-multiple-points>`__ + - `Example box and point input with negative label <#example-box-and-point-input-with-negative-label>`__ + +- `Interactive segmentation <#interactive-segmentation>`__ +- `Run OpenVINO model in automatic mask generation mode <#run-openvino-model-in-automatic-mask-generation-mode>`__ +- `Optimize encoder using NNCF Post-training Quantization API <#optimize-encoder-using-nncf-post-training-quantization-api>`__ + + - `Prepare a calibration dataset <#prepare-a-calibration-dataset>`__ + - `Run quantization and serialize OpenVINO IR model <#run-quantization-and-serialize-openvino-ir-model>`__ + - `Validate Quantized Model Inference <#validate-quantized-model-inference>`__ + - `Compare Performance of the Original and Quantized Models <#compare-performance-of-the-original-and-quantized-models>`__ + Segmentation - identifying which image pixels belong to an object - is a core task in computer vision and is used in a broad array of applications, from analyzing scientific imagery to editing photos. But @@ -25,8 +55,9 @@ zero-shot transfer). This notebook shows an example of how to convert and use Segment Anything Model in OpenVINO format, allowing it to run on a variety of platforms that support an OpenVINO. -Background ----------- +Background `⇑ <#top>`__ +############################################################################################################################### + Previously, to solve any kind of segmentation problem, there were two classes of approaches. The first, interactive segmentation, allowed for @@ -93,24 +124,35 @@ post `__ +############################################################################################################################### + .. code:: ipython3 !pip install -q "segment_anything" "gradio>=3.25" -Convert model to OpenVINO Intermediate Representation ------------------------------------------------------ -Download model checkpoint and create PyTorch model -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +.. parsed-literal:: + + + [notice] A new release of pip is available: 23.1.2 -> 23.2 + [notice] To update, run: pip install --upgrade pip + + +Convert model to OpenVINO Intermediate Representation `⇑ <#top>`__ +############################################################################################################################### + + +Download model checkpoint and create PyTorch model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + There are several Segment Anything Model `checkpoints `__ available for downloading In this tutorial we will use model based on ``vit_b``, but the demonstrated approach is very general and applicable -to other SAM models. Set the model url, path for saving checkpoint and +to other SAM models. Set the model URL, path for saving checkpoint and model type below to a SAM model checkpoint, then load the model using ``sam_model_registry``. @@ -154,8 +196,9 @@ into account this fact, we split model on 2 independent parts: image_encoder and mask_predictor (combination of Prompt Encoder and Mask Decoder). -Image Encoder -~~~~~~~~~~~~~ +Image Encoder `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Image Encoder input is tensor with shape ``1x3x1024x1024`` in ``NCHW`` format, contains image for segmentation. Image Encoder output is image @@ -185,10 +228,36 @@ embeddings, tensor with shape ``1x256x64x64`` serialize(ov_encoder_model, str(ov_encoder_path)) else: ov_encoder_model = core.read_model(ov_encoder_path) - ov_encoder = core.compile_model(ov_encoder_model) -Mask predictor -~~~~~~~~~~~~~~ +.. code:: ipython3 + + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=2, options=('CPU', 'GPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + ov_encoder = core.compile_model(ov_encoder_model, device.value) + +Mask predictor `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + This notebook expects the model was exported with the parameter ``return_single_mask=True``. It means that model will only return the @@ -198,16 +267,27 @@ images this can improve runtime when upscaling masks is expensive. Combined prompt encoder and mask decoder model has following list of inputs: -* ``image_embeddings``: The image embedding from ``image_encoder``. Has a batch index of length 1. -* ``point_coords``: Coordinates of sparse input prompts, corresponding to both point inputs and box inputs. Boxes are encoded using two points, one for the top-left corner and one for the bottom-right corner. *Coordinates must already be transformed to long-side 1024.* Has a batch index of length 1. -* ``point_labels``: Labels for the sparse input prompts. 0 is a negative input point, 1 is a positive input point, 2 is a top-left box corner, 3 is a bottom-right box corner, and -1 is a padding point. -* If there is no box input, a single padding point with label -1 and coordinates (0.0,0.0) should be concatenated. +- ``image_embeddings``: The image embedding from ``image_encoder``. Has + a batch index of length 1. +- ``point_coords``: Coordinates of sparse input prompts, corresponding + to both point inputs and box inputs. Boxes are encoded using two + points, one for the top-left corner and one for the bottom-right + corner. *Coordinates must already be transformed to long-side 1024.* + Has a batch index of length 1. +- ``point_labels``: Labels for the sparse input prompts. 0 is a + negative input point, 1 is a positive input point, 2 is a top-left + box corner, 3 is a bottom-right box corner, and -1 is a padding + point. \*If there is no box input, a single padding point with label + -1 and coordinates (0.0, 0.0) should be concatenated. -Model outputs: +Model outputs: -* ``masks`` - predicted masks resized to original image size, to obtain a binary mask, should be compared with ``threshold`` (usually equal 0.0). -* ``iou_predictions`` - intersection over union predictions -* ``low_res_masks`` - predicted masks before postprocessing, can be used as mask input for model. +- ``masks`` - predicted masks resized to original image size, to obtain + a binary mask, should be compared with ``threshold`` (usually equal + 0.0). +- ``iou_predictions`` - intersection over union predictions +- ``low_res_masks`` - predicted masks before postprocessing, can be + used as mask input for model. .. code:: ipython3 @@ -353,13 +433,31 @@ Model outputs: serialize(ov_model, str(ov_model_path)) else: ov_model = core.read_model(ov_model_path) - ov_predictor = core.compile_model(ov_model) -Run OpenVINO model in interactive segmentation mode ---------------------------------------------------- +.. code:: ipython3 + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=2, options=('CPU', 'GPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + ov_predictor = core.compile_model(ov_model, device.value) + +Run OpenVINO model in interactive segmentation mode `⇑ <#top>`__ +############################################################################################################################### + + +Example Image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Example Image -~~~~~~~~~~~~~ .. code:: ipython3 @@ -386,21 +484,22 @@ Example Image -.. image:: 237-segment-anything-with-output_files/237-segment-anything-with-output_17_0.png +.. image:: 237-segment-anything-with-output_files/237-segment-anything-with-output_21_0.png -Preprocessing and visualization utilities -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Preprocessing and visualization utilities `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -To prepare iinput for Image Encoder we should: + +To prepare input for Image Encoder we should: 1. Convert BGR image to RGB 2. Resize image saving aspect ratio where longest size equal to Image Encoder input size - 1024. 3. Normalize image subtract mean values (123.675, 116.28, 103.53) and divide by std (58.395, 57.12, 57.375) -4. transpose HWC data layout to CHW and add batch dimension. -5. add zero padding to input tensor by height or width (depends on +4. Transpose HWC data layout to CHW and add batch dimension. +5. Add zero padding to input tensor by height or width (depends on aspect ratio) according Image Encoder expected input shape. These steps are applicable to all available models @@ -505,8 +604,9 @@ These steps are applicable to all available models w, h = box[2] - box[0], box[3] - box[1] ax.add_patch(plt.Rectangle((x0, y0), w, h, edgecolor='green', facecolor=(0, 0, 0, 0), lw=2)) -Image encoding -~~~~~~~~~~~~~~ +Image encoding `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + To start work with image, we should preprocess it and obtain image embeddings using ``ov_encoder``. We will use the same image for all @@ -522,8 +622,9 @@ reuse them. Now, we can try to provide different prompts for mask generation -Example point input -~~~~~~~~~~~~~~~~~~~ +Example point input `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + In this example we select one point. The green star symbol show its location on the image below. @@ -541,7 +642,7 @@ location on the image below. -.. image:: 237-segment-anything-with-output_files/237-segment-anything-with-output_24_0.png +.. image:: 237-segment-anything-with-output_files/237-segment-anything-with-output_28_0.png Add a batch index, concatenate a padding point, and transform it to @@ -585,11 +686,12 @@ object). -.. image:: 237-segment-anything-with-output_files/237-segment-anything-with-output_31_0.png +.. image:: 237-segment-anything-with-output_files/237-segment-anything-with-output_35_0.png -Example with multiple points -~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Example with multiple points `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + in this example, we provide additional point for cover larger object area. @@ -611,7 +713,7 @@ Now, prompt for model looks like represented on this image: -.. image:: 237-segment-anything-with-output_files/237-segment-anything-with-output_35_0.png +.. image:: 237-segment-anything-with-output_files/237-segment-anything-with-output_39_0.png Transform the points as in the previous example. @@ -650,13 +752,14 @@ Package inputs, then predict and threshold the mask. -.. image:: 237-segment-anything-with-output_files/237-segment-anything-with-output_40_0.png +.. image:: 237-segment-anything-with-output_files/237-segment-anything-with-output_44_0.png Great! Looks like now, predicted mask cover whole truck. -Example box and point input with negative label -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Example box and point input with negative label `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + In this example we define input prompt using bounding box and point inside it.The bounding box represented as set of points of its left @@ -680,7 +783,7 @@ point should be excluded from mask. -.. image:: 237-segment-anything-with-output_files/237-segment-anything-with-output_44_0.png +.. image:: 237-segment-anything-with-output_files/237-segment-anything-with-output_48_0.png Add a batch index, concatenate a box and point inputs, add the @@ -725,11 +828,12 @@ Package inputs, then predict and threshold the mask. -.. image:: 237-segment-anything-with-output_files/237-segment-anything-with-output_49_0.png +.. image:: 237-segment-anything-with-output_files/237-segment-anything-with-output_53_0.png -Interactive segmentation ------------------------- +Interactive segmentation `⇑ <#top>`__ +############################################################################################################################### + Now, you can try SAM on own image. Upload image to input window and click on desired point, model predict segment based on your image and @@ -817,7 +921,15 @@ point. .. parsed-literal:: - Running on local URL: http://127.0.0.1:7860 + /tmp/ipykernel_1187339/1907223323.py:46: GradioDeprecationWarning: The `style` method is deprecated. Please set these arguments in the constructor instead. + input_img = gr.Image(label="Input", type="numpy").style(height=480, width=480) + /tmp/ipykernel_1187339/1907223323.py:47: GradioDeprecationWarning: The `style` method is deprecated. Please set these arguments in the constructor instead. + output_img = gr.Image(label="Selected Segment", type="numpy").style(height=480, width=480) + + +.. parsed-literal:: + + Running on local URL: http://127.0.0.1:7862 To create a public link, set `share=True` in `launch()`. @@ -825,11 +937,12 @@ point. .. raw:: html -
+
-Run OpenVINO model in automatic mask generation mode ----------------------------------------------------- +Run OpenVINO model in automatic mask generation mode `⇑ <#top>`__ +############################################################################################################################### + Since SAM can efficiently process prompts, masks for the entire image can be generated by sampling a large number of prompts over an image. @@ -1136,13 +1249,15 @@ smaller objects, and post-processing can remove stray pixels and holes ``automatic_mask_generation`` returns a list over masks, where each mask is a dictionary containing various data about the mask. These keys are: -* ``segmentation`` : the mask -* ``area`` : the area of the mask in pixels -* ``bbox`` : the boundary box of the mask in XYWH format -* ``predicted_iou`` : the model’s own prediction for the quality of the mask -* ``point_coords`` : the sampled input point that generated this mask -* ``stability_score`` : an additional measure of mask quality -* ``crop_box`` : the crop of the image used to generate this mask in XYWH format +- ``segmentation`` : the mask +- ``area`` : the area of the mask in pixels +- ``bbox`` : the boundary box of the mask in XYWH format +- ``predicted_iou`` : the model’s own prediction for the quality of the + mask +- ``point_coords`` : the sampled input point that generated this mask +- ``stability_score`` : an additional measure of mask quality +- ``crop_box`` : the crop of the image used to generate this mask in + XYWH format .. code:: ipython3 @@ -1189,12 +1304,13 @@ is a dictionary containing various data about the mask. These keys are: -.. image:: 237-segment-anything-with-output_files/237-segment-anything-with-output_64_1.png +.. image:: 237-segment-anything-with-output_files/237-segment-anything-with-output_68_1.png -Optimize encoder using NNCF Post-training Quantization API ----------------------------------------------------------- +Optimize encoder using NNCF Post-training Quantization API `⇑ <#top>`__ +############################################################################################################################### + `NNCF `__ provides a suite of advanced algorithms for Neural Networks inference optimization in @@ -1211,8 +1327,9 @@ The optimization process contains the following steps: 3. Serialize OpenVINO IR model, using the ``openvino.runtime.serialize`` function. -Prepare a calibration dataset -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Prepare a calibration dataset `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Download COCO dataset. Since the dataset is used to calibrate the model’s parameter instead of fine-tuning it, we don’t need to download @@ -1283,21 +1400,14 @@ dataset and returns data that can be passed to the model for inference. calibration_dataset = nncf.Dataset(calibration_loader, transform_fn) -.. parsed-literal:: - - /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/openvino/offline_transformations/__init__.py:10: FutureWarning: The module is private and following namespace `offline_transformations` will be removed in the future. - warnings.warn( - Post-training Optimization Tool is deprecated and will be removed in the future. Please use Neural Network Compression Framework instead: https://github.com/openvinotoolkit/nncf - Nevergrad package could not be imported. If you are planning to use any hyperparameter optimization algo, consider installing it using pip. This implies advanced usage of the tool. Note that nevergrad is compatible only with Python 3.7+ - - .. parsed-literal:: INFO:nncf:NNCF initialized successfully. Supported frameworks detected: torch, tensorflow, onnx, openvino -Run quantization and serialize OpenVINO IR model -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Run quantization and serialize OpenVINO IR model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The ``nncf.quantize`` function provides an interface for model quantization. It requires an instance of the OpenVINO Model and @@ -1325,24 +1435,452 @@ activations. quantized_model = nncf.quantize(model, calibration_dataset, model_type=nncf.parameters.ModelType.TRANSFORMER, - preset=nncf.common.quantization.structs.QuantizationPreset.MIXED) + preset=nncf.common.quantization.structs.QuantizationPreset.MIXED, subset_size=128) print("model quantization finished") .. parsed-literal:: - INFO:openvino.tools.pot.pipeline.pipeline:Inference Engine version: 2023.0.0-10862-40bf400b189 - INFO:openvino.tools.pot.pipeline.pipeline:Model Optimizer version: 2023.0.0-10862-40bf400b189 - INFO:openvino.tools.pot.pipeline.pipeline:Post-Training Optimization Tool version: 2023.0.0-10862-40bf400b189 - INFO:openvino.tools.pot.statistics.collector:Start computing statistics for algorithms : DefaultQuantization - INFO:openvino.tools.pot.statistics.collector:Computing statistics finished - INFO:openvino.tools.pot.pipeline.pipeline:Start algorithm: DefaultQuantization - INFO:openvino.tools.pot.algorithms.quantization.default.algorithm:Start computing statistics for algorithm : ActivationChannelAlignment - INFO:openvino.tools.pot.algorithms.quantization.default.algorithm:Computing statistics finished - INFO:openvino.tools.pot.algorithms.quantization.default.algorithm:Start computing statistics for algorithms : MinMaxQuantization,FastBiasCorrection - INFO:openvino.tools.pot.algorithms.quantization.default.algorithm:Computing statistics finished - INFO:openvino.tools.pot.pipeline.pipeline:Finished: DefaultQuantization - =========================================================================== + INFO:nncf:709 ignored nodes was found by types in the NNCFGraph + INFO:nncf:24 ignored nodes was found by name in the NNCFGraph + INFO:nncf:Not adding activation input quantizer for operation: 6 /Add + INFO:nncf:Not adding activation input quantizer for operation: 9 /blocks.0/norm1/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 10 /blocks.0/norm1/Sub + INFO:nncf:Not adding activation input quantizer for operation: 16 /blocks.0/norm1/Pow + INFO:nncf:Not adding activation input quantizer for operation: 24 /blocks.0/norm1/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 34 /blocks.0/norm1/Add + INFO:nncf:Not adding activation input quantizer for operation: 45 /blocks.0/norm1/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 15 /blocks.0/norm1/Div + INFO:nncf:Not adding activation input quantizer for operation: 23 /blocks.0/norm1/Mul + 33 /blocks.0/norm1/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 556 /blocks.0/attn/Squeeze + INFO:nncf:Not adding activation input quantizer for operation: 557 /blocks.0/attn/Squeeze_1 + INFO:nncf:Not adding activation input quantizer for operation: 558 /blocks.0/attn/Squeeze_2 + INFO:nncf:Not adding activation input quantizer for operation: 633 /blocks.0/attn/Mul_2 + INFO:nncf:Not adding activation input quantizer for operation: 472 /blocks.0/attn/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 552 /blocks.0/attn/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 551 /blocks.0/attn/Softmax + INFO:nncf:Not adding activation input quantizer for operation: 631 /blocks.0/attn/MatMul_1 + INFO:nncf:Not adding activation input quantizer for operation: 8 /blocks.0/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 13 /blocks.0/norm2/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 14 /blocks.0/norm2/Sub + INFO:nncf:Not adding activation input quantizer for operation: 22 /blocks.0/norm2/Pow + INFO:nncf:Not adding activation input quantizer for operation: 32 /blocks.0/norm2/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 43 /blocks.0/norm2/Add + INFO:nncf:Not adding activation input quantizer for operation: 56 /blocks.0/norm2/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 21 /blocks.0/norm2/Div + INFO:nncf:Not adding activation input quantizer for operation: 31 /blocks.0/norm2/Mul + 42 /blocks.0/norm2/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 91 /blocks.0/mlp/act/Div + INFO:nncf:Not adding activation input quantizer for operation: 154 /blocks.0/mlp/act/Add + INFO:nncf:Not adding activation input quantizer for operation: 92 /blocks.0/mlp/act/Mul + INFO:nncf:Not adding activation input quantizer for operation: 120 /blocks.0/mlp/act/Mul_1 + INFO:nncf:Not adding activation input quantizer for operation: 12 /blocks.0/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 19 /blocks.1/norm1/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 20 /blocks.1/norm1/Sub + INFO:nncf:Not adding activation input quantizer for operation: 30 /blocks.1/norm1/Pow + INFO:nncf:Not adding activation input quantizer for operation: 41 /blocks.1/norm1/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 54 /blocks.1/norm1/Add + INFO:nncf:Not adding activation input quantizer for operation: 72 /blocks.1/norm1/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 29 /blocks.1/norm1/Div + INFO:nncf:Not adding activation input quantizer for operation: 40 /blocks.1/norm1/Mul + 53 /blocks.1/norm1/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 731 /blocks.1/attn/Squeeze + INFO:nncf:Not adding activation input quantizer for operation: 732 /blocks.1/attn/Squeeze_1 + INFO:nncf:Not adding activation input quantizer for operation: 733 /blocks.1/attn/Squeeze_2 + INFO:nncf:Not adding activation input quantizer for operation: 820 /blocks.1/attn/Mul_2 + INFO:nncf:Not adding activation input quantizer for operation: 616 /blocks.1/attn/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 727 /blocks.1/attn/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 726 /blocks.1/attn/Softmax + INFO:nncf:Not adding activation input quantizer for operation: 818 /blocks.1/attn/MatMul_1 + INFO:nncf:Not adding activation input quantizer for operation: 18 /blocks.1/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 27 /blocks.1/norm2/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 28 /blocks.1/norm2/Sub + INFO:nncf:Not adding activation input quantizer for operation: 39 /blocks.1/norm2/Pow + INFO:nncf:Not adding activation input quantizer for operation: 52 /blocks.1/norm2/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 66 /blocks.1/norm2/Add + INFO:nncf:Not adding activation input quantizer for operation: 85 /blocks.1/norm2/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 38 /blocks.1/norm2/Div + INFO:nncf:Not adding activation input quantizer for operation: 51 /blocks.1/norm2/Mul + 65 /blocks.1/norm2/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 140 /blocks.1/mlp/act/Div + INFO:nncf:Not adding activation input quantizer for operation: 272 /blocks.1/mlp/act/Add + INFO:nncf:Not adding activation input quantizer for operation: 141 /blocks.1/mlp/act/Mul + INFO:nncf:Not adding activation input quantizer for operation: 201 /blocks.1/mlp/act/Mul_1 + INFO:nncf:Not adding activation input quantizer for operation: 26 /blocks.1/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 36 /blocks.2/norm1/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 37 /blocks.2/norm1/Sub + INFO:nncf:Not adding activation input quantizer for operation: 50 /blocks.2/norm1/Pow + INFO:nncf:Not adding activation input quantizer for operation: 64 /blocks.2/norm1/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 83 /blocks.2/norm1/Add + INFO:nncf:Not adding activation input quantizer for operation: 107 /blocks.2/norm1/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 49 /blocks.2/norm1/Div + INFO:nncf:Not adding activation input quantizer for operation: 63 /blocks.2/norm1/Mul + 82 /blocks.2/norm1/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 525 /blocks.2/attn/Squeeze + INFO:nncf:Not adding activation input quantizer for operation: 526 /blocks.2/attn/Squeeze_1 + INFO:nncf:Not adding activation input quantizer for operation: 527 /blocks.2/attn/Squeeze_2 + INFO:nncf:Not adding activation input quantizer for operation: 605 /blocks.2/attn/Mul_2 + INFO:nncf:Not adding activation input quantizer for operation: 436 /blocks.2/attn/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 521 /blocks.2/attn/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 520 /blocks.2/attn/Softmax + INFO:nncf:Not adding activation input quantizer for operation: 603 /blocks.2/attn/MatMul_1 + INFO:nncf:Not adding activation input quantizer for operation: 35 /blocks.2/Add + INFO:nncf:Not adding activation input quantizer for operation: 47 /blocks.2/norm2/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 48 /blocks.2/norm2/Sub + INFO:nncf:Not adding activation input quantizer for operation: 62 /blocks.2/norm2/Pow + INFO:nncf:Not adding activation input quantizer for operation: 81 /blocks.2/norm2/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 102 /blocks.2/norm2/Add + INFO:nncf:Not adding activation input quantizer for operation: 135 /blocks.2/norm2/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 61 /blocks.2/norm2/Div + INFO:nncf:Not adding activation input quantizer for operation: 80 /blocks.2/norm2/Mul + 101 /blocks.2/norm2/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 253 /blocks.2/mlp/act/Div + INFO:nncf:Not adding activation input quantizer for operation: 427 /blocks.2/mlp/act/Add + INFO:nncf:Not adding activation input quantizer for operation: 254 /blocks.2/mlp/act/Mul + INFO:nncf:Not adding activation input quantizer for operation: 330 /blocks.2/mlp/act/Mul_1 + INFO:nncf:Not adding activation input quantizer for operation: 46 /blocks.2/Add_1 + INFO:nncf:Not adding activation input quantizer for operation: 59 /blocks.3/norm1/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 60 /blocks.3/norm1/Sub + INFO:nncf:Not adding activation input quantizer for operation: 79 /blocks.3/norm1/Pow + INFO:nncf:Not adding activation input quantizer for operation: 100 /blocks.3/norm1/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 133 /blocks.3/norm1/Add + INFO:nncf:Not adding activation input quantizer for operation: 174 /blocks.3/norm1/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 78 /blocks.3/norm1/Div + INFO:nncf:Not adding activation input quantizer for operation: 99 /blocks.3/norm1/Mul + 132 /blocks.3/norm1/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 1110 /blocks.3/attn/Squeeze + INFO:nncf:Not adding activation input quantizer for operation: 1111 /blocks.3/attn/Squeeze_1 + INFO:nncf:Not adding activation input quantizer for operation: 1112 /blocks.3/attn/Squeeze_2 + INFO:nncf:Not adding activation input quantizer for operation: 1192 /blocks.3/attn/Mul_2 + INFO:nncf:Not adding activation input quantizer for operation: 1013 /blocks.3/attn/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 1106 /blocks.3/attn/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 1105 /blocks.3/attn/Softmax + INFO:nncf:Not adding activation input quantizer for operation: 1190 /blocks.3/attn/MatMul_1 + INFO:nncf:Not adding activation input quantizer for operation: 58 /blocks.3/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 76 /blocks.3/norm2/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 77 /blocks.3/norm2/Sub + INFO:nncf:Not adding activation input quantizer for operation: 98 /blocks.3/norm2/Pow + INFO:nncf:Not adding activation input quantizer for operation: 131 /blocks.3/norm2/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 168 /blocks.3/norm2/Add + INFO:nncf:Not adding activation input quantizer for operation: 247 /blocks.3/norm2/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 97 /blocks.3/norm2/Div + INFO:nncf:Not adding activation input quantizer for operation: 130 /blocks.3/norm2/Mul + 167 /blocks.3/norm2/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 413 /blocks.3/mlp/act/Div + INFO:nncf:Not adding activation input quantizer for operation: 588 /blocks.3/mlp/act/Add + INFO:nncf:Not adding activation input quantizer for operation: 414 /blocks.3/mlp/act/Mul + INFO:nncf:Not adding activation input quantizer for operation: 506 /blocks.3/mlp/act/Mul_1 + INFO:nncf:Not adding activation input quantizer for operation: 75 /blocks.3/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 95 /blocks.4/norm1/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 96 /blocks.4/norm1/Sub + INFO:nncf:Not adding activation input quantizer for operation: 129 /blocks.4/norm1/Pow + INFO:nncf:Not adding activation input quantizer for operation: 166 /blocks.4/norm1/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 245 /blocks.4/norm1/Add + INFO:nncf:Not adding activation input quantizer for operation: 317 /blocks.4/norm1/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 128 /blocks.4/norm1/Div + INFO:nncf:Not adding activation input quantizer for operation: 165 /blocks.4/norm1/Mul + 244 /blocks.4/norm1/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 1294 /blocks.4/attn/Squeeze + INFO:nncf:Not adding activation input quantizer for operation: 1295 /blocks.4/attn/Squeeze_1 + INFO:nncf:Not adding activation input quantizer for operation: 1296 /blocks.4/attn/Squeeze_2 + INFO:nncf:Not adding activation input quantizer for operation: 1384 /blocks.4/attn/Mul_2 + INFO:nncf:Not adding activation input quantizer for operation: 1176 /blocks.4/attn/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 1290 /blocks.4/attn/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 1289 /blocks.4/attn/Softmax + INFO:nncf:Not adding activation input quantizer for operation: 1382 /blocks.4/attn/MatMul_1 + INFO:nncf:Not adding activation input quantizer for operation: 94 /blocks.4/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 126 /blocks.4/norm2/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 127 /blocks.4/norm2/Sub + INFO:nncf:Not adding activation input quantizer for operation: 164 /blocks.4/norm2/Pow + INFO:nncf:Not adding activation input quantizer for operation: 243 /blocks.4/norm2/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 311 /blocks.4/norm2/Add + INFO:nncf:Not adding activation input quantizer for operation: 407 /blocks.4/norm2/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 163 /blocks.4/norm2/Div + INFO:nncf:Not adding activation input quantizer for operation: 242 /blocks.4/norm2/Mul + 310 /blocks.4/norm2/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 574 /blocks.4/mlp/act/Div + INFO:nncf:Not adding activation input quantizer for operation: 777 /blocks.4/mlp/act/Add + INFO:nncf:Not adding activation input quantizer for operation: 575 /blocks.4/mlp/act/Mul + INFO:nncf:Not adding activation input quantizer for operation: 678 /blocks.4/mlp/act/Mul_1 + INFO:nncf:Not adding activation input quantizer for operation: 125 /blocks.4/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 161 /blocks.5/norm1/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 162 /blocks.5/norm1/Sub + INFO:nncf:Not adding activation input quantizer for operation: 241 /blocks.5/norm1/Pow + INFO:nncf:Not adding activation input quantizer for operation: 309 /blocks.5/norm1/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 405 /blocks.5/norm1/Add + INFO:nncf:Not adding activation input quantizer for operation: 493 /blocks.5/norm1/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 240 /blocks.5/norm1/Div + INFO:nncf:Not adding activation input quantizer for operation: 308 /blocks.5/norm1/Mul + 404 /blocks.5/norm1/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 1079 /blocks.5/attn/Squeeze + INFO:nncf:Not adding activation input quantizer for operation: 1080 /blocks.5/attn/Squeeze_1 + INFO:nncf:Not adding activation input quantizer for operation: 1081 /blocks.5/attn/Squeeze_2 + INFO:nncf:Not adding activation input quantizer for operation: 1165 /blocks.5/attn/Mul_2 + INFO:nncf:Not adding activation input quantizer for operation: 977 /blocks.5/attn/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 1075 /blocks.5/attn/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 1074 /blocks.5/attn/Softmax + INFO:nncf:Not adding activation input quantizer for operation: 1163 /blocks.5/attn/MatMul_1 + INFO:nncf:Not adding activation input quantizer for operation: 160 /blocks.5/Add + INFO:nncf:Not adding activation input quantizer for operation: 238 /blocks.5/norm2/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 239 /blocks.5/norm2/Sub + INFO:nncf:Not adding activation input quantizer for operation: 307 /blocks.5/norm2/Pow + INFO:nncf:Not adding activation input quantizer for operation: 403 /blocks.5/norm2/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 488 /blocks.5/norm2/Add + INFO:nncf:Not adding activation input quantizer for operation: 569 /blocks.5/norm2/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 306 /blocks.5/norm2/Div + INFO:nncf:Not adding activation input quantizer for operation: 402 /blocks.5/norm2/Mul + 487 /blocks.5/norm2/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 758 /blocks.5/mlp/act/Div + INFO:nncf:Not adding activation input quantizer for operation: 968 /blocks.5/mlp/act/Add + INFO:nncf:Not adding activation input quantizer for operation: 759 /blocks.5/mlp/act/Mul + INFO:nncf:Not adding activation input quantizer for operation: 859 /blocks.5/mlp/act/Mul_1 + INFO:nncf:Not adding activation input quantizer for operation: 237 /blocks.5/Add_1 + INFO:nncf:Not adding activation input quantizer for operation: 304 /blocks.6/norm1/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 305 /blocks.6/norm1/Sub + INFO:nncf:Not adding activation input quantizer for operation: 401 /blocks.6/norm1/Pow + INFO:nncf:Not adding activation input quantizer for operation: 486 /blocks.6/norm1/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 567 /blocks.6/norm1/Add + INFO:nncf:Not adding activation input quantizer for operation: 651 /blocks.6/norm1/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 400 /blocks.6/norm1/Div + INFO:nncf:Not adding activation input quantizer for operation: 485 /blocks.6/norm1/Mul + 566 /blocks.6/norm1/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 1661 /blocks.6/attn/Squeeze + INFO:nncf:Not adding activation input quantizer for operation: 1662 /blocks.6/attn/Squeeze_1 + INFO:nncf:Not adding activation input quantizer for operation: 1663 /blocks.6/attn/Squeeze_2 + INFO:nncf:Not adding activation input quantizer for operation: 1734 /blocks.6/attn/Mul_2 + INFO:nncf:Not adding activation input quantizer for operation: 1571 /blocks.6/attn/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 1657 /blocks.6/attn/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 1656 /blocks.6/attn/Softmax + INFO:nncf:Not adding activation input quantizer for operation: 1732 /blocks.6/attn/MatMul_1 + INFO:nncf:Not adding activation input quantizer for operation: 303 /blocks.6/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 398 /blocks.6/norm2/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 399 /blocks.6/norm2/Sub + INFO:nncf:Not adding activation input quantizer for operation: 484 /blocks.6/norm2/Pow + INFO:nncf:Not adding activation input quantizer for operation: 565 /blocks.6/norm2/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 645 /blocks.6/norm2/Add + INFO:nncf:Not adding activation input quantizer for operation: 752 /blocks.6/norm2/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 483 /blocks.6/norm2/Div + INFO:nncf:Not adding activation input quantizer for operation: 564 /blocks.6/norm2/Mul + 644 /blocks.6/norm2/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 954 /blocks.6/mlp/act/Div + INFO:nncf:Not adding activation input quantizer for operation: 1148 /blocks.6/mlp/act/Add + INFO:nncf:Not adding activation input quantizer for operation: 955 /blocks.6/mlp/act/Mul + INFO:nncf:Not adding activation input quantizer for operation: 1060 /blocks.6/mlp/act/Mul_1 + INFO:nncf:Not adding activation input quantizer for operation: 397 /blocks.6/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 481 /blocks.7/norm1/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 482 /blocks.7/norm1/Sub + INFO:nncf:Not adding activation input quantizer for operation: 563 /blocks.7/norm1/Pow + INFO:nncf:Not adding activation input quantizer for operation: 643 /blocks.7/norm1/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 750 /blocks.7/norm1/Add + INFO:nncf:Not adding activation input quantizer for operation: 846 /blocks.7/norm1/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 562 /blocks.7/norm1/Div + INFO:nncf:Not adding activation input quantizer for operation: 642 /blocks.7/norm1/Mul + 749 /blocks.7/norm1/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 1821 /blocks.7/attn/Squeeze + INFO:nncf:Not adding activation input quantizer for operation: 1822 /blocks.7/attn/Squeeze_1 + INFO:nncf:Not adding activation input quantizer for operation: 1823 /blocks.7/attn/Squeeze_2 + INFO:nncf:Not adding activation input quantizer for operation: 1897 /blocks.7/attn/Mul_2 + INFO:nncf:Not adding activation input quantizer for operation: 1718 /blocks.7/attn/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 1817 /blocks.7/attn/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 1816 /blocks.7/attn/Softmax + INFO:nncf:Not adding activation input quantizer for operation: 1895 /blocks.7/attn/MatMul_1 + INFO:nncf:Not adding activation input quantizer for operation: 480 /blocks.7/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 560 /blocks.7/norm2/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 561 /blocks.7/norm2/Sub + INFO:nncf:Not adding activation input quantizer for operation: 641 /blocks.7/norm2/Pow + INFO:nncf:Not adding activation input quantizer for operation: 748 /blocks.7/norm2/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 840 /blocks.7/norm2/Add + INFO:nncf:Not adding activation input quantizer for operation: 948 /blocks.7/norm2/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 640 /blocks.7/norm2/Div + INFO:nncf:Not adding activation input quantizer for operation: 747 /blocks.7/norm2/Mul + 839 /blocks.7/norm2/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 1134 /blocks.7/mlp/act/Div + INFO:nncf:Not adding activation input quantizer for operation: 1341 /blocks.7/mlp/act/Add + INFO:nncf:Not adding activation input quantizer for operation: 1135 /blocks.7/mlp/act/Mul + INFO:nncf:Not adding activation input quantizer for operation: 1241 /blocks.7/mlp/act/Mul_1 + INFO:nncf:Not adding activation input quantizer for operation: 559 /blocks.7/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 638 /blocks.8/norm1/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 639 /blocks.8/norm1/Sub + INFO:nncf:Not adding activation input quantizer for operation: 746 /blocks.8/norm1/Pow + INFO:nncf:Not adding activation input quantizer for operation: 838 /blocks.8/norm1/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 946 /blocks.8/norm1/Add + INFO:nncf:Not adding activation input quantizer for operation: 1047 /blocks.8/norm1/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 745 /blocks.8/norm1/Div + INFO:nncf:Not adding activation input quantizer for operation: 837 /blocks.8/norm1/Mul + 945 /blocks.8/norm1/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 1630 /blocks.8/attn/Squeeze + INFO:nncf:Not adding activation input quantizer for operation: 1631 /blocks.8/attn/Squeeze_1 + INFO:nncf:Not adding activation input quantizer for operation: 1632 /blocks.8/attn/Squeeze_2 + INFO:nncf:Not adding activation input quantizer for operation: 1707 /blocks.8/attn/Mul_2 + INFO:nncf:Not adding activation input quantizer for operation: 1535 /blocks.8/attn/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 1626 /blocks.8/attn/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 1625 /blocks.8/attn/Softmax + INFO:nncf:Not adding activation input quantizer for operation: 1705 /blocks.8/attn/MatMul_1 + INFO:nncf:Not adding activation input quantizer for operation: 637 /blocks.8/Add + INFO:nncf:Not adding activation input quantizer for operation: 743 /blocks.8/norm2/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 744 /blocks.8/norm2/Sub + INFO:nncf:Not adding activation input quantizer for operation: 836 /blocks.8/norm2/Pow + INFO:nncf:Not adding activation input quantizer for operation: 944 /blocks.8/norm2/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 1042 /blocks.8/norm2/Add + INFO:nncf:Not adding activation input quantizer for operation: 1129 /blocks.8/norm2/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 835 /blocks.8/norm2/Div + INFO:nncf:Not adding activation input quantizer for operation: 943 /blocks.8/norm2/Mul + 1041 /blocks.8/norm2/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 1322 /blocks.8/mlp/act/Div + INFO:nncf:Not adding activation input quantizer for operation: 1526 /blocks.8/mlp/act/Add + INFO:nncf:Not adding activation input quantizer for operation: 1323 /blocks.8/mlp/act/Mul + INFO:nncf:Not adding activation input quantizer for operation: 1422 /blocks.8/mlp/act/Mul_1 + INFO:nncf:Not adding activation input quantizer for operation: 742 /blocks.8/Add_1 + INFO:nncf:Not adding activation input quantizer for operation: 833 /blocks.9/norm1/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 834 /blocks.9/norm1/Sub + INFO:nncf:Not adding activation input quantizer for operation: 942 /blocks.9/norm1/Pow + INFO:nncf:Not adding activation input quantizer for operation: 1040 /blocks.9/norm1/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 1127 /blocks.9/norm1/Add + INFO:nncf:Not adding activation input quantizer for operation: 1214 /blocks.9/norm1/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 941 /blocks.9/norm1/Div + INFO:nncf:Not adding activation input quantizer for operation: 1039 /blocks.9/norm1/Mul + 1126 /blocks.9/norm1/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 2098 /blocks.9/attn/Squeeze + INFO:nncf:Not adding activation input quantizer for operation: 2099 /blocks.9/attn/Squeeze_1 + INFO:nncf:Not adding activation input quantizer for operation: 2100 /blocks.9/attn/Squeeze_2 + INFO:nncf:Not adding activation input quantizer for operation: 2137 /blocks.9/attn/Mul_2 + INFO:nncf:Not adding activation input quantizer for operation: 2038 /blocks.9/attn/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 2094 /blocks.9/attn/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 2093 /blocks.9/attn/Softmax + INFO:nncf:Not adding activation input quantizer for operation: 2135 /blocks.9/attn/MatMul_1 + INFO:nncf:Not adding activation input quantizer for operation: 832 /blocks.9/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 939 /blocks.9/norm2/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 940 /blocks.9/norm2/Sub + INFO:nncf:Not adding activation input quantizer for operation: 1038 /blocks.9/norm2/Pow + INFO:nncf:Not adding activation input quantizer for operation: 1125 /blocks.9/norm2/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 1208 /blocks.9/norm2/Add + INFO:nncf:Not adding activation input quantizer for operation: 1316 /blocks.9/norm2/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 1037 /blocks.9/norm2/Div + INFO:nncf:Not adding activation input quantizer for operation: 1124 /blocks.9/norm2/Mul + 1207 /blocks.9/norm2/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 1512 /blocks.9/mlp/act/Div + INFO:nncf:Not adding activation input quantizer for operation: 1690 /blocks.9/mlp/act/Add + INFO:nncf:Not adding activation input quantizer for operation: 1513 /blocks.9/mlp/act/Mul + INFO:nncf:Not adding activation input quantizer for operation: 1611 /blocks.9/mlp/act/Mul_1 + INFO:nncf:Not adding activation input quantizer for operation: 938 /blocks.9/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 1035 /blocks.10/norm1/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 1036 /blocks.10/norm1/Sub + INFO:nncf:Not adding activation input quantizer for operation: 1123 /blocks.10/norm1/Pow + INFO:nncf:Not adding activation input quantizer for operation: 1206 /blocks.10/norm1/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 1314 /blocks.10/norm1/Add + INFO:nncf:Not adding activation input quantizer for operation: 1409 /blocks.10/norm1/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 1122 /blocks.10/norm1/Div + INFO:nncf:Not adding activation input quantizer for operation: 1205 /blocks.10/norm1/Mul + 1313 /blocks.10/norm1/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 2155 /blocks.10/attn/Squeeze + INFO:nncf:Not adding activation input quantizer for operation: 2156 /blocks.10/attn/Squeeze_1 + INFO:nncf:Not adding activation input quantizer for operation: 2157 /blocks.10/attn/Squeeze_2 + INFO:nncf:Not adding activation input quantizer for operation: 2177 /blocks.10/attn/Mul_2 + INFO:nncf:Not adding activation input quantizer for operation: 2121 /blocks.10/attn/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 2151 /blocks.10/attn/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 2150 /blocks.10/attn/Softmax + INFO:nncf:Not adding activation input quantizer for operation: 2175 /blocks.10/attn/MatMul_1 + INFO:nncf:Not adding activation input quantizer for operation: 1034 /blocks.10/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 1120 /blocks.10/norm2/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 1121 /blocks.10/norm2/Sub + INFO:nncf:Not adding activation input quantizer for operation: 1204 /blocks.10/norm2/Pow + INFO:nncf:Not adding activation input quantizer for operation: 1312 /blocks.10/norm2/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 1403 /blocks.10/norm2/Add + INFO:nncf:Not adding activation input quantizer for operation: 1506 /blocks.10/norm2/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 1203 /blocks.10/norm2/Div + INFO:nncf:Not adding activation input quantizer for operation: 1311 /blocks.10/norm2/Mul + 1402 /blocks.10/norm2/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 1676 /blocks.10/mlp/act/Div + INFO:nncf:Not adding activation input quantizer for operation: 1854 /blocks.10/mlp/act/Add + INFO:nncf:Not adding activation input quantizer for operation: 1677 /blocks.10/mlp/act/Mul + INFO:nncf:Not adding activation input quantizer for operation: 1768 /blocks.10/mlp/act/Mul_1 + INFO:nncf:Not adding activation input quantizer for operation: 1119 /blocks.10/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 1201 /blocks.11/norm1/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 1202 /blocks.11/norm1/Sub + INFO:nncf:Not adding activation input quantizer for operation: 1310 /blocks.11/norm1/Pow + INFO:nncf:Not adding activation input quantizer for operation: 1401 /blocks.11/norm1/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 1504 /blocks.11/norm1/Add + INFO:nncf:Not adding activation input quantizer for operation: 1598 /blocks.11/norm1/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 1309 /blocks.11/norm1/Div + INFO:nncf:Not adding activation input quantizer for operation: 1400 /blocks.11/norm1/Mul + 1503 /blocks.11/norm1/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 2067 /blocks.11/attn/Squeeze + INFO:nncf:Not adding activation input quantizer for operation: 2068 /blocks.11/attn/Squeeze_1 + INFO:nncf:Not adding activation input quantizer for operation: 2069 /blocks.11/attn/Squeeze_2 + INFO:nncf:Not adding activation input quantizer for operation: 2110 /blocks.11/attn/Mul_2 + INFO:nncf:Not adding activation input quantizer for operation: 2002 /blocks.11/attn/Add_2 + INFO:nncf:Not adding activation input quantizer for operation: 2063 /blocks.11/attn/Add_3 + INFO:nncf:Not adding activation input quantizer for operation: 2062 /blocks.11/attn/Softmax + INFO:nncf:Not adding activation input quantizer for operation: 2108 /blocks.11/attn/MatMul_1 + INFO:nncf:Not adding activation input quantizer for operation: 1200 /blocks.11/Add + INFO:nncf:Not adding activation input quantizer for operation: 1307 /blocks.11/norm2/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 1308 /blocks.11/norm2/Sub + INFO:nncf:Not adding activation input quantizer for operation: 1399 /blocks.11/norm2/Pow + INFO:nncf:Not adding activation input quantizer for operation: 1502 /blocks.11/norm2/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 1593 /blocks.11/norm2/Add + INFO:nncf:Not adding activation input quantizer for operation: 1671 /blocks.11/norm2/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 1398 /blocks.11/norm2/Div + INFO:nncf:Not adding activation input quantizer for operation: 1501 /blocks.11/norm2/Mul + 1592 /blocks.11/norm2/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 1835 /blocks.11/mlp/act/Div + INFO:nncf:Not adding activation input quantizer for operation: 1993 /blocks.11/mlp/act/Add + INFO:nncf:Not adding activation input quantizer for operation: 1836 /blocks.11/mlp/act/Mul + INFO:nncf:Not adding activation input quantizer for operation: 1913 /blocks.11/mlp/act/Mul_1 + INFO:nncf:Not adding activation input quantizer for operation: 1306 /blocks.11/Add_1 + INFO:nncf:Not adding activation input quantizer for operation: 1590 /neck/neck.1/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 1591 /neck/neck.1/Sub + INFO:nncf:Not adding activation input quantizer for operation: 1669 /neck/neck.1/Pow + INFO:nncf:Not adding activation input quantizer for operation: 1741 /neck/neck.1/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 1834 /neck/neck.1/Add + INFO:nncf:Not adding activation input quantizer for operation: 1911 /neck/neck.1/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 1668 /neck/neck.1/Div + INFO:nncf:Not adding activation input quantizer for operation: 1740 /neck/neck.1/Mul + 1833 /neck/neck.1/Add_1 + + INFO:nncf:Not adding activation input quantizer for operation: 1991 /neck/neck.3/ReduceMean + INFO:nncf:Not adding activation input quantizer for operation: 1992 /neck/neck.3/Sub + INFO:nncf:Not adding activation input quantizer for operation: 2058 /neck/neck.3/Pow + INFO:nncf:Not adding activation input quantizer for operation: 2106 /neck/neck.3/ReduceMean_1 + INFO:nncf:Not adding activation input quantizer for operation: 2144 /neck/neck.3/Add + INFO:nncf:Not adding activation input quantizer for operation: 2168 /neck/neck.3/Sqrt + INFO:nncf:Not adding activation input quantizer for operation: 2057 /neck/neck.3/Div + INFO:nncf:Not adding activation input quantizer for operation: 2105 /neck/neck.3/Mul + 2143 4017 + + + +.. parsed-literal:: + + Statistics collection: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 128/128 [05:14<00:00, 2.45s/it] + Biases correction: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 48/48 [06:34<00:00, 8.21s/it] + +.. parsed-literal:: + model quantization finished @@ -1351,8 +1889,9 @@ activations. ov_encoder_path_int8 = "sam_image_encoder_int8.xml" serialize(quantized_model, ov_encoder_path_int8) -Validate Quantized Model Inference -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Validate Quantized Model Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + We can reuse the previous code to validate the output of ``INT8`` model. @@ -1360,7 +1899,7 @@ We can reuse the previous code to validate the output of ``INT8`` model. # Load INT8 model and run pipeline again ov_encoder_model_int8 = core.read_model(ov_encoder_path_int8) - ov_encoder_int8 = core.compile_model(ov_encoder_model_int8) + ov_encoder_int8 = core.compile_model(ov_encoder_model_int8, device.value) encoding_results = ov_encoder_int8(preprocessed_image) image_embeddings = encoding_results[ov_encoder_int8.output(0)] @@ -1389,7 +1928,7 @@ We can reuse the previous code to validate the output of ``INT8`` model. -.. image:: 237-segment-anything-with-output_files/237-segment-anything-with-output_76_0.png +.. image:: 237-segment-anything-with-output_files/237-segment-anything-with-output_80_0.png Run ``INT8`` model in automatic mask generation mode @@ -1406,17 +1945,17 @@ Run ``INT8`` model in automatic mask generation mode .. parsed-literal:: - 0%| | 0/46 [00:00`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ Finally, use the OpenVINO `Benchmark Tool `__ @@ -1426,7 +1965,7 @@ models. .. code:: ipython3 # Inference FP32 model (OpenVINO IR) - !benchmark_app -m $ov_encoder_path -d 'CPU' + !benchmark_app -m $ov_encoder_path -d $device.value .. parsed-literal:: @@ -1434,19 +1973,20 @@ models. [Step 1/11] Parsing and validating input arguments [ INFO ] Parsing input parameters [Step 2/11] Loading OpenVINO Runtime + [ WARNING ] Default duration 120 seconds is used for unknown device AUTO [ INFO ] OpenVINO: - [ INFO ] Build ................................. 2023.0.0-10862-40bf400b189 + [ INFO ] Build ................................. 2023.0.1-11005-fa1c41994f3-releases/2023/0 [ INFO ] [ INFO ] Device info: - [ INFO ] CPU - [ INFO ] Build ................................. 2023.0.0-10862-40bf400b189 + [ INFO ] AUTO + [ INFO ] Build ................................. 2023.0.1-11005-fa1c41994f3-releases/2023/0 [ INFO ] [ INFO ] [Step 3/11] Setting device configuration - [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.THROUGHPUT. + [ WARNING ] Performance hint was not explicitly specified in command line. Device(AUTO) performance hint will be set to PerformanceMode.THROUGHPUT. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 103.19 ms + [ INFO ] Read model took 69.37 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] input.1 (node: input.1) : f32 / [...] / [1,3,1024,1024] @@ -1460,45 +2000,52 @@ models. [ INFO ] Model outputs: [ INFO ] 4017 (node: 4017) : f32 / [...] / [1,256,64,64] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 1085.86 ms + [ INFO ] Compile model took 1196.87 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: + [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT [ INFO ] NETWORK_NAME: torch_jit [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 12 - [ INFO ] NUM_STREAMS: 12 - [ INFO ] AFFINITY: Affinity.CORE - [ INFO ] INFERENCE_NUM_THREADS: 36 - [ INFO ] PERF_COUNT: False - [ INFO ] INFERENCE_PRECISION_HINT: - [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT - [ INFO ] EXECUTION_MODE_HINT: ExecutionMode.PERFORMANCE - [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 - [ INFO ] ENABLE_CPU_PINNING: True - [ INFO ] SCHEDULING_CORE_TYPE: SchedulingCoreType.ANY_CORE - [ INFO ] ENABLE_HYPER_THREADING: True + [ INFO ] MODEL_PRIORITY: Priority.MEDIUM + [ INFO ] MULTI_DEVICE_PRIORITIES: CPU + [ INFO ] CPU: + [ INFO ] CPU_BIND_THREAD: YES + [ INFO ] CPU_THREADS_NUM: 0 + [ INFO ] CPU_THROUGHPUT_STREAMS: 12 + [ INFO ] DEVICE_ID: + [ INFO ] DUMP_EXEC_GRAPH_AS_DOT: + [ INFO ] DYN_BATCH_ENABLED: NO + [ INFO ] DYN_BATCH_LIMIT: 0 + [ INFO ] ENFORCE_BF16: NO + [ INFO ] EXCLUSIVE_ASYNC_REQUESTS: NO + [ INFO ] NETWORK_NAME: torch_jit + [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 12 + [ INFO ] PERFORMANCE_HINT: THROUGHPUT + [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 + [ INFO ] PERF_COUNT: NO [ INFO ] EXECUTION_DEVICES: ['CPU'] [Step 9/11] Creating infer requests and preparing input tensors [ WARNING ] No input files were given for input 'input.1'!. This input will be filled with random values! [ INFO ] Fill input 'input.1' with random values - [Step 10/11] Measuring performance (Start inference asynchronously, 12 inference requests, limits: 60000 ms duration) + [Step 10/11] Measuring performance (Start inference asynchronously, 12 inference requests, limits: 120000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 4292.83 ms + [ INFO ] First inference took 4043.51 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 60 iterations - [ INFO ] Duration: 75716.93 ms + [ INFO ] Count: 108 iterations + [ INFO ] Duration: 135037.41 ms [ INFO ] Latency: - [ INFO ] Median: 14832.33 ms - [ INFO ] Average: 14780.77 ms - [ INFO ] Min: 10398.47 ms - [ INFO ] Max: 16725.65 ms - [ INFO ] Throughput: 0.79 FPS + [ INFO ] Median: 14646.89 ms + [ INFO ] Average: 14615.54 ms + [ INFO ] Min: 6295.79 ms + [ INFO ] Max: 19356.55 ms + [ INFO ] Throughput: 0.80 FPS .. code:: ipython3 # Inference INT8 model (OpenVINO IR) - !benchmark_app -m $ov_encoder_path_int8 -d 'CPU' + !benchmark_app -m $ov_encoder_path_int8 -d $device.value .. parsed-literal:: @@ -1506,19 +2053,20 @@ models. [Step 1/11] Parsing and validating input arguments [ INFO ] Parsing input parameters [Step 2/11] Loading OpenVINO Runtime + [ WARNING ] Default duration 120 seconds is used for unknown device AUTO [ INFO ] OpenVINO: - [ INFO ] Build ................................. 2023.0.0-10862-40bf400b189 + [ INFO ] Build ................................. 2023.0.1-11005-fa1c41994f3-releases/2023/0 [ INFO ] [ INFO ] Device info: - [ INFO ] CPU - [ INFO ] Build ................................. 2023.0.0-10862-40bf400b189 + [ INFO ] AUTO + [ INFO ] Build ................................. 2023.0.1-11005-fa1c41994f3-releases/2023/0 [ INFO ] [ INFO ] [Step 3/11] Setting device configuration - [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.THROUGHPUT. + [ WARNING ] Performance hint was not explicitly specified in command line. Device(AUTO) performance hint will be set to PerformanceMode.THROUGHPUT. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 123.84 ms + [ INFO ] Read model took 104.31 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] input.1 (node: input.1) : f32 / [...] / [1,3,1024,1024] @@ -1532,37 +2080,44 @@ models. [ INFO ] Model outputs: [ INFO ] 4017 (node: 4017) : f32 / [...] / [1,256,64,64] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 2132.36 ms + [ INFO ] Compile model took 1414.62 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: + [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT [ INFO ] NETWORK_NAME: torch_jit [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 12 - [ INFO ] NUM_STREAMS: 12 - [ INFO ] AFFINITY: Affinity.CORE - [ INFO ] INFERENCE_NUM_THREADS: 36 - [ INFO ] PERF_COUNT: False - [ INFO ] INFERENCE_PRECISION_HINT: - [ INFO ] PERFORMANCE_HINT: PerformanceMode.THROUGHPUT - [ INFO ] EXECUTION_MODE_HINT: ExecutionMode.PERFORMANCE - [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 - [ INFO ] ENABLE_CPU_PINNING: True - [ INFO ] SCHEDULING_CORE_TYPE: SchedulingCoreType.ANY_CORE - [ INFO ] ENABLE_HYPER_THREADING: True + [ INFO ] MODEL_PRIORITY: Priority.MEDIUM + [ INFO ] MULTI_DEVICE_PRIORITIES: CPU + [ INFO ] CPU: + [ INFO ] CPU_BIND_THREAD: YES + [ INFO ] CPU_THREADS_NUM: 0 + [ INFO ] CPU_THROUGHPUT_STREAMS: 12 + [ INFO ] DEVICE_ID: + [ INFO ] DUMP_EXEC_GRAPH_AS_DOT: + [ INFO ] DYN_BATCH_ENABLED: NO + [ INFO ] DYN_BATCH_LIMIT: 0 + [ INFO ] ENFORCE_BF16: NO + [ INFO ] EXCLUSIVE_ASYNC_REQUESTS: NO + [ INFO ] NETWORK_NAME: torch_jit + [ INFO ] OPTIMAL_NUMBER_OF_INFER_REQUESTS: 12 + [ INFO ] PERFORMANCE_HINT: THROUGHPUT + [ INFO ] PERFORMANCE_HINT_NUM_REQUESTS: 0 + [ INFO ] PERF_COUNT: NO [ INFO ] EXECUTION_DEVICES: ['CPU'] [Step 9/11] Creating infer requests and preparing input tensors [ WARNING ] No input files were given for input 'input.1'!. This input will be filled with random values! [ INFO ] Fill input 'input.1' with random values - [Step 10/11] Measuring performance (Start inference asynchronously, 12 inference requests, limits: 60000 ms duration) + [Step 10/11] Measuring performance (Start inference asynchronously, 12 inference requests, limits: 120000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 3113.94 ms + [ INFO ] First inference took 2694.03 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 72 iterations - [ INFO ] Duration: 68936.14 ms + [ INFO ] Count: 132 iterations + [ INFO ] Duration: 129404.57 ms [ INFO ] Latency: - [ INFO ] Median: 11281.87 ms - [ INFO ] Average: 11162.87 ms - [ INFO ] Min: 6736.09 ms - [ INFO ] Max: 12547.48 ms - [ INFO ] Throughput: 1.04 FPS + [ INFO ] Median: 11651.20 ms + [ INFO ] Average: 11526.49 ms + [ INFO ] Min: 5003.59 ms + [ INFO ] Max: 13329.53 ms + [ INFO ] Throughput: 1.02 FPS diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_17_0.png b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_17_0.png deleted file mode 100644 index 723651f629b..00000000000 --- a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_17_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:25049f23b031c74a21254b4fefabae004dfcdbb5a168aa9c676e26d79b0e016f -size 467418 diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_21_0.png b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_21_0.png new file mode 100644 index 00000000000..c6bb0b4bdca --- /dev/null +++ b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_21_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:300b3c9db63a5fdb80cf4156f6c50ac72d64359111791ed9dfca167071a3d8e8 +size 467418 diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_24_0.png b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_24_0.png deleted file mode 100644 index 2a431805152..00000000000 --- a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_24_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:f1f59b5c414758c7490ce9f609ad091af6115b76838060dfe62e7b338ec27373 -size 468529 diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_28_0.png b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_28_0.png new file mode 100644 index 00000000000..a4484cc8c5a --- /dev/null +++ b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_28_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:8d7ad488d386ef255fcc9843a4c52297f9242456be6652945676d27aa66e36bf +size 468529 diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_31_0.png b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_31_0.png deleted file mode 100644 index e3cd5fb4e2d..00000000000 --- a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_31_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:a81039f74801f315b4608146d85a5cc22e86cb1e05609c09158ead64f1c52d38 -size 469443 diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_35_0.png b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_35_0.png index 04e047d7d8b..5b7f897a0fc 100644 --- a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_35_0.png +++ b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_35_0.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:a320efee0b674c3a9cce9c56f0bc17b0855a9792c4f7a2b9d7744f58740b7a43 -size 470668 +oid sha256:2d7bea1506407cac81d630310269f5903c2a5b6106dbbad489fc0f02f5c3452f +size 469443 diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_39_0.png b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_39_0.png new file mode 100644 index 00000000000..0036aeaefc8 --- /dev/null +++ b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_39_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:602034052a3cb68853676e3ec0309ed4554c786b8fd6af2cf4d3c6d75fb05cd3 +size 470668 diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_40_0.png b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_40_0.png deleted file mode 100644 index b0046ab972c..00000000000 --- a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_40_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:967bbf4a0669a3cd0c634f81b43c4b1f4864de47359d4ff8abfbc34f1c4f9150 -size 468092 diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_44_0.png b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_44_0.png index 19aab869e6b..9ff95b0d8ab 100644 --- a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_44_0.png +++ b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_44_0.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:bdafdf87053b7fd9fb2fa72c4123ddce5e59451237722ed81b19d203171cf95e -size 468088 +oid sha256:1bf87c5699a07bc45057b539825b1cfd668792f171b74e0a469e95afbc0b99d7 +size 468091 diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_48_0.png b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_48_0.png new file mode 100644 index 00000000000..ef9be641056 --- /dev/null +++ b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_48_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:8fd1894491acc22b9c7eb7db28eee021e11af3c323dd30c7f93a363e6ba55899 +size 468088 diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_49_0.png b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_49_0.png deleted file mode 100644 index ea58766142d..00000000000 --- a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_49_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:6700875076cc176c9b15da84aade168bb8ff8d44266a57c41479d0b601195d4f -size 472755 diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_53_0.png b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_53_0.png new file mode 100644 index 00000000000..04e80ce21bd --- /dev/null +++ b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_53_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:63cec12807e5b72c324cc14400458ea6415ccc1fabfd232fc3a972975f2e0a7e +size 472754 diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_64_1.png b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_64_1.png deleted file mode 100644 index fa1e26ec285..00000000000 --- a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_64_1.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:8fe250d7ee766610cb448a8770dc0e255523f6c5aa1603c5ba51c8bd6d059934 -size 2234575 diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_68_1.jpg b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_68_1.jpg new file mode 100644 index 00000000000..9128bc05b81 --- /dev/null +++ b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_68_1.jpg @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:c8d9b804802277c0d0481f24aa18680a8ef2ea16a9cfd1ab74f26569433b85b3 +size 261936 diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_68_1.png b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_68_1.png new file mode 100644 index 00000000000..2fbb8cd039a --- /dev/null +++ b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_68_1.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:7898288cc74d40771eb95d54c384d6e886a01ab7e307638c727e4968a9f6ddba +size 2431859 diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_76_0.png b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_76_0.png deleted file mode 100644 index 44389f64971..00000000000 --- a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_76_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:3e99643b9ab3ce40e1653dbf80029c6037e3278eb086b4ef23d96675a1bfc38b -size 469385 diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_78_1.png b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_78_1.png deleted file mode 100644 index 4484ccbcb01..00000000000 --- a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_78_1.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:47ba3c8d49706dfeb1ddbf6038057641a22bd80a99d6d2a96d80e368f69cec0d -size 2277366 diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_80_0.png b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_80_0.png new file mode 100644 index 00000000000..5f6c56307ee --- /dev/null +++ b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_80_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:fd592a78aa28237ff5e3738215368873a50a2af828bb84863a1e55d4e764c91a +size 469423 diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_82_1.jpg b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_82_1.jpg new file mode 100644 index 00000000000..62bb879f1cb --- /dev/null +++ b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_82_1.jpg @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:b85995823f6fad126857e3b3ab6eed4304d22648b37df75343ebd4fa9a303ee8 +size 262791 diff --git a/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_82_1.png b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_82_1.png new file mode 100644 index 00000000000..de392dfacb7 --- /dev/null +++ b/docs/notebooks/237-segment-anything-with-output_files/237-segment-anything-with-output_82_1.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:5e810c37a2dc6cd4b05e40a834c024018dc42d1672d5d22a3bfc7fb8f82b748c +size 2434363 diff --git a/docs/notebooks/237-segment-anything-with-output_files/index.html b/docs/notebooks/237-segment-anything-with-output_files/index.html index 25671cbfdb6..28568c66ba6 100644 --- a/docs/notebooks/237-segment-anything-with-output_files/index.html +++ b/docs/notebooks/237-segment-anything-with-output_files/index.html @@ -1,16 +1,18 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/237-segment-anything-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/237-segment-anything-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/237-segment-anything-with-output_files/


../
-237-segment-anything-with-output_17_0.png          12-Jul-2023 00:11              467418
-237-segment-anything-with-output_24_0.png          12-Jul-2023 00:11              468529
-237-segment-anything-with-output_31_0.png          12-Jul-2023 00:11              469443
-237-segment-anything-with-output_35_0.png          12-Jul-2023 00:11              470668
-237-segment-anything-with-output_40_0.png          12-Jul-2023 00:11              468092
-237-segment-anything-with-output_44_0.png          12-Jul-2023 00:11              468088
-237-segment-anything-with-output_49_0.png          12-Jul-2023 00:11              472755
-237-segment-anything-with-output_64_1.png          12-Jul-2023 00:11             2234575
-237-segment-anything-with-output_76_0.png          12-Jul-2023 00:11              469385
-237-segment-anything-with-output_78_1.png          12-Jul-2023 00:11             2277366
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/237-segment-anything-with-output_files/


../
+237-segment-anything-with-output_21_0.png          16-Aug-2023 01:31              467418
+237-segment-anything-with-output_28_0.png          16-Aug-2023 01:31              468529
+237-segment-anything-with-output_35_0.png          16-Aug-2023 01:31              469443
+237-segment-anything-with-output_39_0.png          16-Aug-2023 01:31              470668
+237-segment-anything-with-output_44_0.png          16-Aug-2023 01:31              468091
+237-segment-anything-with-output_48_0.png          16-Aug-2023 01:31              468088
+237-segment-anything-with-output_53_0.png          16-Aug-2023 01:31              472754
+237-segment-anything-with-output_68_1.jpg          16-Aug-2023 01:31              261936
+237-segment-anything-with-output_68_1.png          16-Aug-2023 01:31             2431859
+237-segment-anything-with-output_80_0.png          16-Aug-2023 01:31              469423
+237-segment-anything-with-output_82_1.jpg          16-Aug-2023 01:31              262791
+237-segment-anything-with-output_82_1.png          16-Aug-2023 01:31             2434363
 

diff --git a/docs/notebooks/238-deep-floyd-if-with-output.rst b/docs/notebooks/238-deep-floyd-if-with-output.rst index 62cf3d779fd..7585c074bad 100644 --- a/docs/notebooks/238-deep-floyd-if-with-output.rst +++ b/docs/notebooks/238-deep-floyd-if-with-output.rst @@ -1,6 +1,8 @@ Image generation with DeepFloyd IF and OpenVINO™ ================================================ +.. _top: + DeepFloyd IF is an advanced open-source text-to-image model that delivers remarkable photorealism and language comprehension. DeepFloyd IF consists of a frozen text encoder and three cascaded pixel diffusion @@ -73,16 +75,49 @@ vector in embedded space. 3. Stage 3: Follows the same path as Stage 2 and upscales the image to 1024x1024 pixel resolution. It is not released yet, so we will use a - conventional Super Resolution network to get hir-res results. + conventional Super Resolution network to get hi-res results. + - .. note:: +**Table of contents**: - - *This example requires the download of roughly 27 GB of model checkpoints, which could take some time depending on your internet connection speed. Additionally, the converted models will consume another 27 GB of disk space.* - - *Please be aware that a minimum of 32 GB of RAM is necessary to convert and run inference on the models. There may be instances where the notebook appears to freeze or stop responding.* - - *To access the model checkpoints,you'll need a Hugging Face account. You'll also be prompted to explicitly accept the* `model license `__ +- `Prerequisites <#prerequisites>`__ -Prerequisites -------------- + - `Authentication <#authentication>`__ + +- `DeepFloyd IF in Diffusers library <#deepfloyd-if-in-diffusers-library>`__ +- `Convert models to OpenVINO Intermediate representation (IR) format <#convert-models-to-openvino-intermediate-representation-ir-format>`__ +- `Convert Text Encoder <#convert-text-encoder>`__ +- `Convert the first Pixel Diffusion module’s UNet <#convert-the-first-pixel-diffusion-modules-unet>`__ +- `Convert the second pixel diffusion module <#convert-the-second-pixel-diffusion-module>`__ +- `Prepare Inference pipeline <#prepare-inference-pipeline>`__ +- `Run Text-to-Image generation <#run-text-to-image-generation>`__ + + - `Text Encoder inference <#text-encoder-inference>`__ + - `First Stage diffusion block inference <#first-stage-diffusion-block-inference>`__ + - `Second Stage diffusion block inference <#second-stage-diffusion-block-inference>`__ + - `Third Stage diffusion block <#third-stage-diffusion-block>`__ + - `Upscale the generated image using a Super Resolution network <#upscale-the-generated-image-using-a-super-resolution-network>`__ + + - `Download the Super Resolution model weights <#download-the-super-resolution-model-weights>`__ + - `Reshape the model’s inputs <#reshape-the-models-inputs>`__ + - `Prepare the input images and run the model <#prepare-the-input-images-and-run-the-model>`__ + - `Display the result <#display-the-result>`__ + +.. note:: + + - *This example requires the download of roughly 27 GB of model + checkpoints, which could take some time depending on your internet + connection speed. Additionally, the converted models will consume + another 27 GB of disk space.* + - *Please be aware that a minimum of 32 GB of RAM is necessary to + convert and run inference on the models. There may be instances + where the notebook appears to freeze or stop responding.* + - *To access the model checkpoints, you’ll need a Hugging Face + account. You’ll also be prompted to explicitly accept the*\ `model + license `__\ *.* + +Prerequisites `⇑ <#top>`__ +############################################################################################################################### Install required packages. @@ -92,7 +127,7 @@ Install required packages. !pip install -q --upgrade pip !pip install -q "diffusers>=0.16.1" accelerate transformers safetensors sentencepiece huggingface_hub - !pip install -q --pre --upgrade openvino-dev + !pip install -q "openvino-dev>=2023.0.0" .. code:: ipython3 @@ -120,34 +155,33 @@ Install required packages. .. code:: ipython3 - # Set up target computing device - DEVICE = 'CPU' - checkpoint_variant = 'fp16' model_dtype = torch.float32 ir_input_type = 'f32' compress_to_fp16 = False - models_dir = Path('./models').expanduser() + models_dir = Path('./models') models_dir.mkdir(exist_ok=True) encoder_ir_path = models_dir / 'encoder_ir.xml' first_stage_unet_ir_path = models_dir / 'unet_ir_I.xml' second_stage_unet_ir_path = models_dir / 'unet_ir_II.xml' -Authentication -~~~~~~~~~~~~~~ +Authentication `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -In order to access IF checkpoints, users need to provide an -authentication token. +In order to access IF checkpoints, users need to provide an authentication token. If you already have a token, you can input it into the provided form in the next cell. If not, please proceed according to the following -instructions: +instructions: -1. Make sure to have a `Hugging Face `__ account and be logged in -2. Accept the license on the model card of `DeepFloyd/IF-I-M-v1.0 `__ -3. To generate a token, proceed to `this page `__ +1. Make sure to have a `Hugging Face `__ + account and be logged in +2. Accept the license on the model card of + `DeepFloyd/IF-I-M-v1.0 `__ +3. To generate a token, proceed to `this + page `__ Uncheck the ``Add token as git credential?`` box. @@ -165,14 +199,14 @@ Uncheck the ``Add token as git credential?`` box. VBox(children=(HTML(value='
`__ +############################################################################################################################### To work with IF by DeepFloyd Lab, we will use `Hugging Face Diffusers package `__. Diffusers package -exposes the DiffusionPipeline class, simplifying experiments with +exposes the ``DiffusionPipeline`` class, simplifying experiments with diffusion models. The code below demonstrates how to create a -DiffusionPipeline using IF configs: +``DiffusionPipeline`` using IF configs: .. code:: ipython3 @@ -230,14 +264,14 @@ DiffusionPipeline using IF configs: Wall time: 16.1 s -Convert models to OpenVINO Intermediate representation (IR) format ------------------------------------------------------------------- +Convert models to OpenVINO Intermediate representation (IR) format. `⇑ <#top>`__ +############################################################################################################################### -The OpenVINO Model Optimizer enables direct conversion of PyTorch -models. We will utilize the mo.convert_model method to acquire OpenVINO -IR versions of the models. This requires providing a model object, input -data for model tracing, and other relevant parameters. The -use_legacy_frontend=True parameter instructs the Model Optimizer to +Model conversion API enables direct conversion of PyTorch +models. We will utilize the ``mo.convert_model`` method to acquire +OpenVINO IR versions of the models. This requires providing a model +object, input data for model tracing, and other relevant parameters. The +``use_legacy_frontend=True`` parameter instructs model conversion API to employ the ONNX model format as an intermediate step, as opposed to using the PyTorch JIT compiler, which is not optimal for our situation. @@ -250,10 +284,11 @@ The pipeline consists of three important parts: - A Stage 2 U-Net that takes low resolution output from the previous step and the latent representations to upscale the resulting image. -Let us convert each part +Let us convert each part. + +1. Convert Text Encoder `⇑ <#top>`__ +############################################################################################################################### -1. Convert Text Encoder ------------------------ The text encoder is responsible for converting the input prompt, such as “ultra close-up color photo portrait of rainbow owl with deer horns in @@ -271,12 +306,12 @@ and/or the PyTorch-specific ``example_input`` argument. However, in this case, the ``InputCutInfo`` class was utilized to describe the model input and provide it as the ``input`` argument. Using the ``InputCutInfo`` class offers a framework-agnostic solution and enables -the definition of complex inputs. It allows for specifying the input -name, shape, type, and value within a single argument, providing greater +the definition of complex inputs. It allows specifying the input name, +shape, type, and value within a single argument, providing greater flexibility. -To learn more please refer to `Model Optimizer Python API -docs `__ +To learn more, refer to this +`page `__ .. code:: ipython3 @@ -303,8 +338,9 @@ docs `__ Wall time: 1.37 s -Convert the first Pixel Diffusion module’s UNet ------------------------------------------------ +Convert the first Pixel Diffusion module’s UNet `⇑ <#top>`__ +############################################################################################################################### + U-Net model gradually denoises latent image representation guided by text encoder hidden state. @@ -347,16 +383,17 @@ resolution images. Wall time: 298 ms -Convert the second pixel diffusion module ------------------------------------------ +Convert the second pixel diffusion module `⇑ <#top>`__ +############################################################################################################################### + The second Diffusion module in the cascade generates 256x256 pixel images. The second stage pipeline will use bilinear interpolation to upscale the -64x64 image that was generated in the previopus stage to a higher -256x256 resolution. Then it will denoise the image taking into account -the encoded user prompt. +64x64 image that was generated in the previous stage to a higher 256x256 +resolution. Then it will denoise the image taking into account the +encoded user prompt. .. code:: ipython3 @@ -387,8 +424,9 @@ the encoded user prompt. Wall time: 273 ms -Prepare Inference pipeline --------------------------- +Prepare Inference pipeline `⇑ <#top>`__ +############################################################################################################################### + The original pipeline from the source repository will be reused in this example. In order to achieve this, adapter classes were created to @@ -399,6 +437,24 @@ seamlessly into the pipeline. core = Core() +Select inference device +~~~~~~~~~~~~~~~~~~~~~~~ + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + .. code:: ipython3 class TextEncoder: @@ -527,16 +583,18 @@ seamlessly into the pipeline. result_numpy = result[self.unet_openvino.outputs[0]] return result_tuple(torch.tensor(result_numpy, dtype=self.dtype)) -Run Text-to-Image generation ----------------------------- +Run Text-to-Image generation `⇑ <#top>`__ +############################################################################################################################### + Now, we can set a text prompt for image generation and execute the inference pipeline. Optionally, you can also modify the random generator seed for latent state initialization and adjust the number of images to be generated for the given prompt. -Text Encoder inference -~~~~~~~~~~~~~~~~~~~~~~ +Text Encoder inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -546,7 +604,7 @@ Text Encoder inference negative_prompt = 'blurred unreal uncentered occluded' # Initialize TextEncoder wrapper class - stage_1.text_encoder = TextEncoder(encoder_ir_path, dtype=model_dtype, device=DEVICE) + stage_1.text_encoder = TextEncoder(encoder_ir_path, dtype=model_dtype, device=device.value) print('The model has been loaded') # Generate text embeddings @@ -582,8 +640,9 @@ Text Encoder inference -First Stage diffusion block inference -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +First Stage diffusion block inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -599,7 +658,7 @@ First Stage diffusion block inference first_stage_unet_ir_path, stage_1_config, dtype=model_dtype, - device=DEVICE + device=device.value ) print('The model has been loaded') @@ -637,12 +696,13 @@ First Stage diffusion block inference -.. image:: 238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_27_3.png +.. image:: 238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_29_3.png -Second Stage diffusion block inference -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Second Stage diffusion block inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -653,7 +713,7 @@ Second Stage diffusion block inference second_stage_unet_ir_path, stage_2_config, dtype=model_dtype, - device=DEVICE + device=device.value ) print('The model has been loaded') @@ -689,18 +749,19 @@ Second Stage diffusion block inference -.. image:: 238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_29_3.png +.. image:: 238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_31_3.png -Third Stage diffusion block -~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Third Stage diffusion block `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -The final block, which upscales images to a higher resolution (1024x1024 -px), has not been released by DeepFloyd yet. Stay tuned! +The final block, which +upscales images to a higher resolution (1024x1024 px), has not been +released by DeepFloyd yet. Stay tuned! -Upscale the generated image using a Super Resolution network -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Upscale the generated image using a Super Resolution network. `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ Though the third stage has not been officially released, we’ll employ the Super Resolution network from `Example @@ -715,8 +776,9 @@ release! # Temporary requirement !pip install -q matplotlib -Download the Super Resolution model weights -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +Download the Super Resolution model weights `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + .. code:: ipython3 @@ -754,13 +816,13 @@ Download the Super Resolution model weights single-image-super-resolution-1032 already downloaded to models -Reshape the model’s inputs -^^^^^^^^^^^^^^^^^^^^^^^^^^ +Reshape the model’s inputs `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- -We need to reshape the inputs for the model. This is necessary because -the IR model was converted with a different target input resolution. The -Second IF stage returns 256x256 pixel images. Using the 4x -SuperResolution model makes our target image size 1024x1024 pixel. +We need to reshape the inputs for the model. This is necessary because the IR model was converted with +a different target input resolution. The Second IF stage returns 256x256 +pixel images. Using the 4x Super Resolution model makes our target image +size 1024x1024 pixel. .. code:: ipython3 @@ -769,10 +831,11 @@ SuperResolution model makes our target image size 1024x1024 pixel. 0: [1, 3, 256, 256], 1: [1, 3, 1024, 1024] }) - compiled_model = core.compile_model(model=model, device_name=DEVICE) + compiled_model = core.compile_model(model=model, device_name=device.value) + +Prepare the input images and run the model `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- -Prepare the input images and run the model -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ .. code:: ipython3 @@ -790,8 +853,9 @@ Prepare the input images and run the model [input_image_original, input_image_bicubic] )[compiled_model.output(0)] -Display the result -^^^^^^^^^^^^^^^^^^ +Display the result `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + .. code:: ipython3 @@ -813,6 +877,6 @@ Display the result -.. image:: 238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_39_0.png +.. image:: 238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_41_0.png diff --git a/docs/notebooks/238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_27_3.png b/docs/notebooks/238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_27_3.png deleted file mode 100644 index b1b76dc3db8..00000000000 --- a/docs/notebooks/238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_27_3.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:e80d41d97270f5a468796b6f875e4f43a3c0155337bf1ad0ea413eb86f78c0fe -size 10886 diff --git a/docs/notebooks/238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_29_3.png b/docs/notebooks/238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_29_3.png index bfdeb60f610..b1b76dc3db8 100644 --- a/docs/notebooks/238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_29_3.png +++ b/docs/notebooks/238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_29_3.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:09c6abed08c84b93c2262a50412c09c273e5403b4f0a25ad17d6420deefe990a -size 129672 +oid sha256:e80d41d97270f5a468796b6f875e4f43a3c0155337bf1ad0ea413eb86f78c0fe +size 10886 diff --git a/docs/notebooks/238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_31_3.png b/docs/notebooks/238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_31_3.png new file mode 100644 index 00000000000..bfdeb60f610 --- /dev/null +++ b/docs/notebooks/238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_31_3.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:09c6abed08c84b93c2262a50412c09c273e5403b4f0a25ad17d6420deefe990a +size 129672 diff --git a/docs/notebooks/238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_39_0.png b/docs/notebooks/238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_41_0.png similarity index 100% rename from docs/notebooks/238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_39_0.png rename to docs/notebooks/238-deep-floyd-if-with-output_files/238-deep-floyd-if-with-output_41_0.png diff --git a/docs/notebooks/238-deep-floyd-if-with-output_files/index.html b/docs/notebooks/238-deep-floyd-if-with-output_files/index.html index 7b0f16e6f1d..cac17a9f83a 100644 --- a/docs/notebooks/238-deep-floyd-if-with-output_files/index.html +++ b/docs/notebooks/238-deep-floyd-if-with-output_files/index.html @@ -1,9 +1,9 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/238-deep-floyd-if-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/238-deep-floyd-if-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/238-deep-floyd-if-with-output_files/


../
-238-deep-floyd-if-with-output_27_3.png             12-Jul-2023 00:11               10886
-238-deep-floyd-if-with-output_29_3.png             12-Jul-2023 00:11              129672
-238-deep-floyd-if-with-output_39_0.png             12-Jul-2023 00:11             1349174
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/238-deep-floyd-if-with-output_files/


../
+238-deep-floyd-if-with-output_29_3.png             16-Aug-2023 01:31               10886
+238-deep-floyd-if-with-output_31_3.png             16-Aug-2023 01:31              129672
+238-deep-floyd-if-with-output_41_0.png             16-Aug-2023 01:31             1349174
 

diff --git a/docs/notebooks/239-image-bind-with-output.rst b/docs/notebooks/239-image-bind-convert-with-output.rst similarity index 99% rename from docs/notebooks/239-image-bind-with-output.rst rename to docs/notebooks/239-image-bind-convert-with-output.rst index e09ac0c7a22..a4407207b50 100644 --- a/docs/notebooks/239-image-bind-with-output.rst +++ b/docs/notebooks/239-image-bind-convert-with-output.rst @@ -1,6 +1,8 @@ Binding multimodal data using ImageBind and OpenVINO ==================================================== +.. _top: + Exploring the surrounding world, people get information using multiple senses, for example, seeing a busy street and hearing the sounds of car engines. ImageBind introduces an approach that brings machines one step @@ -27,7 +29,8 @@ The tutorial consists of following steps: 1. Download the pre-trained model. 2. Prepare input data examples. -3. Convert the model to OpenVINO Intermediate Representation format (IR). +3. Convert the model to OpenVINO Intermediate Representation format + (IR). 4. Run model inference and analyze results. About ImageBind @@ -55,10 +58,10 @@ post `__. Like all embedding models, there are many potential use cases for ImageBind, among them information retrieval, zero-shot classification, and usage created by ImageBind representation as input for downstream -tasks (e.g. image generation). Some of the potental use-cases +tasks (e.g. image generation). Some of the potential use-cases represented on the image below: -.. figure:: https://scontent-dub4-1.xx.fbcdn.net/v/t39.2365-6/344857966_3220413401583307_2893213081653042531_n.png?_nc_cat=111&ccb=1-7&_nc_sid=ad8a9d&_nc_ohc=anLCQJ70ALEAX8Txj_O&_nc_ht=scontent-dub4-1.xx&oh=00_AfDhDiRx-kE1Tk5ewS7Fgyp6VT2Sr7di267heBQsEetiLQ&oe=647211F7 +.. figure:: https://user-images.githubusercontent.com/29454499/256303836-c8e7b311-0b7b-407c-8610-fd8a803e4197.png :alt: usecases usecases @@ -66,8 +69,26 @@ represented on the image below: In this tutorial, we consider how to use ImageBind for multimodal zero-shot classification. -Prerequisites -------------- +**Table of contents**: + +- `Prerequisites <#prerequisites>`__ +- `Instantiate PyTorch model <#instantiate-pytorch-model>`__ +- `Prepare input data <#prepare-input-data>`__ +- `Convert Model to OpenVINO Intermediate Representation (IR) format <#convert-model-to-openvino-intermediate-representation-ir-format>`__ + + - `Select inference device <#select-inference-device>`__ + +- `Zero-shot classification using ImageBind and OpenVINO <#zero-shot-classification-using-imagebind-and-openvino>`__ + + - `Text-Image classification <#text-image-classification>`__ + - `Text-Audio classification <#text-audio-classification>`__ + - `Image-Audio classification <#image-audio-classification>`__ + +- `Next Steps <#next-steps>`__ + +Prerequisites `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -85,6 +106,14 @@ Prerequisites else: !pip install -q "torchaudio==0.13.1+cpu" --find-links https://download.pytorch.org/whl/torch_stable.html + +.. parsed-literal:: + + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + + .. code:: ipython3 from pathlib import Path @@ -99,11 +128,19 @@ Prerequisites .. parsed-literal:: - /home/ea/work/openvino_notebooks/notebooks/240-image-bind/ImageBind + Cloning into 'ImageBind'... + remote: Enumerating objects: 112, done. + remote: Counting objects: 100% (60/60), done. + remote: Compressing objects: 100% (26/26), done. + remote: Total 112 (delta 43), reused 34 (delta 34), pack-reused 52 + Receiving objects: 100% (112/112), 2.64 MiB | 3.81 MiB/s, done. + Resolving deltas: 100% (50/50), done. + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/239-image-bind/ImageBind -Instantiate PyTorch model -------------------------- +Instantiate PyTorch model `⇑ <#top>`__ +############################################################################################################################### + To start work with the model, we should instantiate the PyTorch model class. ``imagebind_model.imagebind_huge(pretrained=True)`` downloads @@ -118,10 +155,10 @@ card `__. .. code:: ipython3 - import data + import imagebind.data as data import torch - from models import imagebind_model - from models.imagebind_model import ModalityType + from imagebind.models import imagebind_model + from imagebind.models.imagebind_model import ModalityType # Instantiate model model = imagebind_model.imagebind_huge(pretrained=True) @@ -130,22 +167,41 @@ card `__. .. parsed-literal:: - /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/torchvision/transforms/_functional_video.py:6: UserWarning: The 'torchvision.transforms._functional_video' module is deprecated since 0.12 and will be removed in the future. Please use the 'torchvision.transforms.functional' module instead. + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torchvision/transforms/_functional_video.py:6: UserWarning: The 'torchvision.transforms._functional_video' module is deprecated since 0.12 and will be removed in the future. Please use the 'torchvision.transforms.functional' module instead. warnings.warn( - /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/torchvision/transforms/_transforms_video.py:22: UserWarning: The 'torchvision.transforms._transforms_video' module is deprecated since 0.12 and will be removed in the future. Please use the 'torchvision.transforms' module instead. + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torchvision/transforms/_transforms_video.py:22: UserWarning: The 'torchvision.transforms._transforms_video' module is deprecated since 0.12 and will be removed in the future. Please use the 'torchvision.transforms' module instead. warnings.warn( -Prepare input data ------------------- +.. parsed-literal:: + + Downloading imagebind weights to .checkpoints/imagebind_huge.pth ... + + + +.. parsed-literal:: + + 0%| | 0.00/4.47G [00:00`__ +############################################################################################################################### + ImageBind works with data across 6 different modalities. Each of them -requires its steps for preprocessing. ``data`` module is repsonsible for +requires its steps for preprocessing. ``data`` module is responsible for data reading and preprocessing for each modality. -* ``data.load_and_transform_text`` accepts a list of text labels and tokenizes them. -* ``data.load_and_transform_vision_data`` accepts paths to input images, reads them, resizes to save aspect ratio with smaller side size 224, performs center crop, and normalizes data into [0, 1] floating point range. -* ``data.load_and_transofrm_audio_data`` reads audio files from provided paths, splits it on samples, and computes `mel `__ spectrogram. +- ``data.load_and_transform_text`` accepts a list of text labels and + tokenizes them. +- ``data.load_and_transform_vision_data`` accepts paths to input + images, reads them, resizes to save aspect ratio with smaller side + size 224, performs center crop, and normalizes data into [0, 1] + floating point range. +- ``data.load_and_transofrm_audio_data`` reads audio files from + provided paths, splits it on samples, and computes + `mel `__ + spectrogram. .. code:: ipython3 @@ -161,8 +217,8 @@ data reading and preprocessing for each modality. ModalityType.AUDIO: data.load_and_transform_audio_data(audio_paths, "cpu"), } -Convert Model to OpenVINO Intermediate Representation (IR) format ------------------------------------------------------------------ +Convert Model to OpenVINO Intermediate Representation (IR) format. `⇑ <#top>`__ +############################################################################################################################### OpenVINO supports PyTorch through export to the ONNX format. You will use the ``torch.onnx.export`` function for obtaining the ONNX model. You @@ -175,12 +231,12 @@ example, input and output names or dynamic shapes). While ONNX models are directly supported by OpenVINO™ runtime, it can be useful to convert them to IR format to take advantage of advanced -OpenVINO optimization tools and features. You will use `OpenVINO Model -Optimizer Python -API `__ -for conversion model to IR format and compression weights to ``FP16`` -format. ``mo.convert_model`` function returns OpenVINO Model class -instance ready to load on a device or save on disk for next loading. +OpenVINO optimization tools and features. You will use `model conversion +Python +API `__ +to convert model to IR format and compress weights to ``FP16`` format. +The ``mo.convert_model`` function returns OpenVINO Model class instance +ready to load on a device or save on a disk for next loading. ImageBind accepts data that represents different modalities simultaneously in any combinations, however, their processing is @@ -206,8 +262,37 @@ embeddings. from openvino.runtime import serialize, Core core = Core() - device = "CPU" + +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + ov_modality_models = {} modalities = [ModalityType.TEXT, ModalityType.VISION, ModalityType.AUDIO] @@ -227,10 +312,30 @@ embeddings. serialize(ov_model, str(ir_path)) else: ov_model = core.read_model(ir_path) - ov_modality_models[modality] = core.compile_model(ov_model, device) + ov_modality_models[modality] = core.compile_model(ov_model, device.value) + + +.. parsed-literal:: + + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/symbolic_opset9.py:5408: UserWarning: Exporting aten::index operator of advanced indexing in opset 14 is achieved by combination of multiple ONNX operators, including Reshape, Transpose, Concat, and Gather. If indices include negative values, the exported graph will produce incorrect results. + warnings.warn( + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/_internal/jit_utils.py:258: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + _C._jit_pass_onnx_node_shape_type_inference(node, params_dict, opset_version) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + _C._jit_pass_onnx_graph_shape_type_inference( + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + _C._jit_pass_onnx_graph_shape_type_inference( + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/239-image-bind/ImageBind/imagebind/models/multimodal_preprocessors.py:433: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + if x.shape[self.time_dim] == 1: + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/239-image-bind/ImageBind/imagebind/models/multimodal_preprocessors.py:259: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + assert tokens.shape[2] == self.embed_dim + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/239-image-bind/ImageBind/imagebind/models/multimodal_preprocessors.py:74: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + if npatch_per_img == N: + + +Zero-shot classification using ImageBind and OpenVINO `⇑ <#top>`__ +############################################################################################################################### -Zero-shot classification using ImageBind and OpenVINO ------------------------------------------------------ In zero-shot classification, a piece of data is embedded and fed to the model to retrieve a label that corresponds with the contents of the @@ -246,9 +351,13 @@ classification. To perform zero-shot classification using ImageBind we should perform the following steps: -1. Preprocess data batch for requested modalities (one modality in our case treated as a data source, other - as a label). +1. Preprocess data batch for requested modalities (one modality in our + case treated as a data source, other - as a label). 2. Calculate embeddings for each modality. -3. Find dot-product between embeddings vectors to get probabilities matrix. 4. Obtain the label with the highest probability for mapping the source into label space. +3. Find dot-product between embeddings vectors to get probabilities + matrix. +4. Obtain the label with the highest probability for mapping the source + into label space. We already preprocessed data in previous step, now, we should run model inference for getting embeddings. @@ -286,8 +395,9 @@ they represent the same object. image_list = [img.split('/')[-1] for img in image_paths] audio_list = [audio.split('/')[-1] for audio in audio_paths] -Text-Image classification -~~~~~~~~~~~~~~~~~~~~~~~~~ +Text-Image classification `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -297,11 +407,12 @@ Text-Image classification -.. image:: 239-image-bind-with-output_files/239-image-bind-with-output_16_0.png +.. image:: 239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_20_0.png -Text-Audio classification -~~~~~~~~~~~~~~~~~~~~~~~~~ +Text-Audio classification `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -311,11 +422,12 @@ Text-Audio classification -.. image:: 239-image-bind-with-output_files/239-image-bind-with-output_18_0.png +.. image:: 239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_22_0.png -Image-Audio classification -~~~~~~~~~~~~~~~~~~~~~~~~~~ +Image-Audio classification `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -325,7 +437,7 @@ Image-Audio classification -.. image:: 239-image-bind-with-output_files/239-image-bind-with-output_20_0.png +.. image:: 239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_24_0.png Putting all together, we can match text, image, and sound for our data. @@ -349,7 +461,7 @@ Putting all together, we can match text, image, and sound for our data. -.. image:: 239-image-bind-with-output_files/239-image-bind-with-output_22_1.png +.. image:: 239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_26_1.png @@ -380,7 +492,7 @@ Putting all together, we can match text, image, and sound for our data. -.. image:: 239-image-bind-with-output_files/239-image-bind-with-output_23_1.png +.. image:: 239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_27_1.png @@ -411,7 +523,7 @@ Putting all together, we can match text, image, and sound for our data. -.. image:: 239-image-bind-with-output_files/239-image-bind-with-output_24_1.png +.. image:: 239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_28_1.png @@ -426,3 +538,10 @@ Putting all together, we can match text, image, and sound for our data. + +Next Steps `⇑ <#top>`__ +############################################################################################################################### + +Open the `239-image-bind-quantize <239-image-bind-quantize.ipynb>`__ notebook to +quantize the IR model with the Post-training Quantization API of NNCF +and compare ``FP16`` and ``INT8`` models. diff --git a/docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_16_0.png b/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_20_0.png similarity index 100% rename from docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_16_0.png rename to docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_20_0.png diff --git a/docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_18_0.png b/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_22_0.png similarity index 100% rename from docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_18_0.png rename to docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_22_0.png diff --git a/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_24_0.png b/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_24_0.png new file mode 100644 index 00000000000..7e80bd97a6e --- /dev/null +++ b/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_24_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:81ff26fa579513e912e907ecec52388d40e407c707e973e863ba0bcedccaf033 +size 18633 diff --git a/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_26_1.jpg b/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_26_1.jpg new file mode 100644 index 00000000000..57b27fdf471 --- /dev/null +++ b/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_26_1.jpg @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:0afb9c967000fda34e9bc86710e52806c4534a4adebcf7ee4da278c51d42b5e7 +size 36700 diff --git a/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_26_1.png b/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_26_1.png new file mode 100644 index 00000000000..dab9cdf3fbf --- /dev/null +++ b/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_26_1.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:7d41a1f38efda8f946697a868f8bf5a5e839b6826b0d0ed55617adfc18bad923 +size 341289 diff --git a/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_27_1.jpg b/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_27_1.jpg new file mode 100644 index 00000000000..74de64eec53 --- /dev/null +++ b/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_27_1.jpg @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:ec7c92e7c0df9f866fb55f2bad06ce27e6f0caec4b1d1ebe8582381db7bf2b2a +size 71448 diff --git a/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_27_1.png b/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_27_1.png new file mode 100644 index 00000000000..ddf371f75b3 --- /dev/null +++ b/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_27_1.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:0056fb57206d58454fe7b19f74ad0f93cb8b91b00cc19c537809e38e9cdaa8fd +size 839471 diff --git a/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_28_1.jpg b/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_28_1.jpg new file mode 100644 index 00000000000..917fa58d82a --- /dev/null +++ b/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_28_1.jpg @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:13bcb163019054d5667867f976db4c158b18ee3bc0d9c385794ff8461e55888f +size 54208 diff --git a/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_28_1.png b/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_28_1.png new file mode 100644 index 00000000000..a149b46c668 --- /dev/null +++ b/docs/notebooks/239-image-bind-convert-with-output_files/239-image-bind-convert-with-output_28_1.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:93a8fe24712027cf386d66d8aac8bb0fc3d208f45b9e27f33d762bfb5d62693a +size 658748 diff --git a/docs/notebooks/239-image-bind-convert-with-output_files/index.html b/docs/notebooks/239-image-bind-convert-with-output_files/index.html new file mode 100644 index 00000000000..651c9506015 --- /dev/null +++ b/docs/notebooks/239-image-bind-convert-with-output_files/index.html @@ -0,0 +1,15 @@ + +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/239-image-bind-convert-with-output_files/ + +

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/239-image-bind-convert-with-output_files/


../
+239-image-bind-convert-with-output_20_0.png        16-Aug-2023 01:31               15385
+239-image-bind-convert-with-output_22_0.png        16-Aug-2023 01:31               13795
+239-image-bind-convert-with-output_24_0.png        16-Aug-2023 01:31               18633
+239-image-bind-convert-with-output_26_1.jpg        16-Aug-2023 01:31               36700
+239-image-bind-convert-with-output_26_1.png        16-Aug-2023 01:31              341289
+239-image-bind-convert-with-output_27_1.jpg        16-Aug-2023 01:31               71448
+239-image-bind-convert-with-output_27_1.png        16-Aug-2023 01:31              839471
+239-image-bind-convert-with-output_28_1.jpg        16-Aug-2023 01:31               54208
+239-image-bind-convert-with-output_28_1.png        16-Aug-2023 01:31              658748
+

+ diff --git a/docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_20_0.png b/docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_20_0.png deleted file mode 100644 index a45a8c5ef50..00000000000 --- a/docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_20_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:f7873673b4249603105fed4e6d89d844f065bffd2f70f3e8166ad7058673620f -size 18151 diff --git a/docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_22_1.png b/docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_22_1.png deleted file mode 100644 index 6be4611dbc7..00000000000 --- a/docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_22_1.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:d09352f8474421fa78d601cc5afbe88df3d0403c157f91605d424b66a2f1809a -size 303014 diff --git a/docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_23_1.png b/docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_23_1.png deleted file mode 100644 index 174dcfdcbe8..00000000000 --- a/docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_23_1.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:609e506939d69a89fb59d36622d72005d5b162afccf70c1e2463cd51d544d4dd -size 777583 diff --git a/docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_24_1.png b/docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_24_1.png deleted file mode 100644 index a4b0b02a4d7..00000000000 --- a/docs/notebooks/239-image-bind-with-output_files/239-image-bind-with-output_24_1.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:7509d532217e990ed721424c57aecbadfb634d397bd1c069852f873fee8741a9 -size 572170 diff --git a/docs/notebooks/239-image-bind-with-output_files/index.html b/docs/notebooks/239-image-bind-with-output_files/index.html deleted file mode 100644 index 41e6bfb50ae..00000000000 --- a/docs/notebooks/239-image-bind-with-output_files/index.html +++ /dev/null @@ -1,12 +0,0 @@ - -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/239-image-bind-with-output_files/ - -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/239-image-bind-with-output_files/


../
-239-image-bind-with-output_16_0.png                12-Jul-2023 00:11               15385
-239-image-bind-with-output_18_0.png                12-Jul-2023 00:11               13795
-239-image-bind-with-output_20_0.png                12-Jul-2023 00:11               18151
-239-image-bind-with-output_22_1.png                12-Jul-2023 00:11              303014
-239-image-bind-with-output_23_1.png                12-Jul-2023 00:11              777583
-239-image-bind-with-output_24_1.png                12-Jul-2023 00:11              572170
-

- diff --git a/docs/notebooks/240-dolly-2-instruction-following-with-output.rst b/docs/notebooks/240-dolly-2-instruction-following-with-output.rst index 2dc1b12409a..ff759a53d68 100644 --- a/docs/notebooks/240-dolly-2-instruction-following-with-output.rst +++ b/docs/notebooks/240-dolly-2-instruction-following-with-output.rst @@ -1,6 +1,8 @@ Instruction following using Databricks Dolly 2.0 and OpenVINO ============================================================= +.. _top: + The instruction following is one of the cornerstones of the current generation of large language models(LLMs). Reinforcement learning with human preferences (`RLHF `__) and @@ -79,8 +81,27 @@ dataset can be found in `Databricks blog post `__ and `repo `__ -Prerequisites -------------- + +**Table of contents**: + +- `Prerequisites <#prerequisites>`__ + + - `Select inference device <#select-inference-device>`__ + +- `Download and Convert Model <#download-and-convert-model>`__ +- `Create an instruction-following inference pipeline <#create-an-instruction-following-inference-pipeline>`__ + + - `Setup imports <#setup-imports>`__ + - `Prepare template for user prompt <#prepare-template-for-user-prompt>`__ + - `Helpers for output parsing <#helpers-for-output-parsing>`__ + - `Main generation function <#main-generation-function>`__ + - `Helpers for application <#helpers-for-application>`__ + +- `Run instruction-following pipeline <#run-instruction-following-pipeline>`__ + +Prerequisites `⇑ <#top>`__ +############################################################################################################################### + First, we should install the `Hugging Face Optimum `__ library @@ -95,15 +116,59 @@ documentation `__. !pip install -q "diffusers>=0.16.1" "transformers>=4.28.0" !pip install -q "git+https://github.com/huggingface/optimum-intel.git" datasets onnx onnxruntime gradio -Download and Convert Model --------------------------- + +.. parsed-literal:: + + + [notice] A new release of pip is available: 23.1.2 -> 23.2 + [notice] To update, run: pip install --upgrade pip + + [notice] A new release of pip is available: 23.1.2 -> 23.2 + [notice] To update, run: pip install --upgrade pip + + +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + from openvino.runtime import Core + + core = Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=2, options=('CPU', 'GPU', 'AUTO'), value='AUTO') + + + +Download and Convert Model `⇑ <#top>`__ +############################################################################################################################### + Optimum Intel can be used to load optimized models from the `Hugging Face Hub `__ and create pipelines to run an inference with OpenVINO Runtime using Hugging Face APIs. The Optimum Inference models are API compatible with Hugging Face Transformers models. This means we just need to replace -AutoModelForXxx class with the corresponding OVModelForXxx class. +``AutoModelForXxx`` class with the corresponding ``OVModelForXxx`` +class. Below is an example of the Dolly model @@ -134,7 +199,7 @@ Tokenizer class and pipelines API are compatible with Optimum models. tokenizer = AutoTokenizer.from_pretrained(model_id) - current_device = "CPU" + current_device = device.value if model_path.exists(): ov_model = OVModelForCausalLM.from_pretrained(model_path, device=current_device) @@ -145,14 +210,10 @@ Tokenizer class and pipelines API are compatible with Optimum models. .. parsed-literal:: - 2023-06-01 12:40:22.260551: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. - 2023-06-01 12:40:22.297683: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + 2023-07-17 14:47:00.308996: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-07-17 14:47:00.348466: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. - 2023-06-01 12:40:22.937097: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT - /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/openvino/offline_transformations/__init__.py:10: FutureWarning: The module is private and following namespace `offline_transformations` will be removed in the future. - warnings.warn( - Post-training Optimization Tool is deprecated and will be removed in the future. Please use Neural Network Compression Framework instead: https://github.com/openvinotoolkit/nncf - Nevergrad package could not be imported. If you are planning to use any hyperparameter optimization algo, consider installing it using pip. This implies advanced usage of the tool. Note that nevergrad is compatible only with Python 3.7+ + 2023-07-17 14:47:01.039895: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT .. parsed-literal:: @@ -160,15 +221,61 @@ Tokenizer class and pipelines API are compatible with Optimum models. INFO:nncf:NNCF initialized successfully. Supported frameworks detected: torch, tensorflow, onnx, openvino -.. parsed-literal:: +.. code:: No CUDA runtime is found, using CUDA_HOME='/usr/local/cuda' comet_ml is installed but `COMET_API_KEY` is not set. - Compiling the model and creating the inference request ... + The argument `from_transformers` is deprecated, and will be removed in optimum 2.0. Use `export` instead + Framework not specified. Using pt to export to ONNX. + Using framework PyTorch: 1.13.1+cpu + Overriding 1 configuration item(s) + - use_cache -> True + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/transformers/models/gpt_neox/modeling_gpt_neox.py:504: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + assert batch_size > 0, "batch_size has to be defined and > 0" + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/transformers/models/gpt_neox/modeling_gpt_neox.py:270: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + if seq_len > self.max_seq_len_cached: + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/nncf/torch/dynamic_graph/wrappers.py:74: TracerWarning: torch.tensor results are registered as constants in the trace. You can safely ignore this warning if you use this function to create tensors out of constant variables that would be the same every time you call this function. In any other case, this might cause the trace to be incorrect. + op1 = operator(*args, **kwargs) + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + In-place op on output of tensor.shape. See https://pytorch.org/docs/master/onnx.html#avoid-inplace-operations-when-using-tensor-shape-in-tracing-mode + Saving external data to one file... + Compiling the model... + Set CACHE_DIR to /tmp/tmpndw8_20n/model_cache -Create an instruction-following inference pipeline --------------------------------------------------- +Create an instruction-following inference pipeline `⇑ <#top>`__ +############################################################################################################################### + The ``run_generation`` function accepts user-provided text input, tokenizes it, and runs the generation process. Text generation is an @@ -198,22 +305,30 @@ find more information about the most popular decoding methods in this There are several parameters that can control text generation quality: -- ``Temperature`` is a parameter used to control the level of creativity in AI-generated text. By adjusting the ``temperature``, you can influence the AI model’s probability distribution, making the text more focused or diverse. +- | ``Temperature`` is a parameter used to control the level of + creativity in AI-generated text. By adjusting the ``temperature``, + you can influence the AI model’s probability distribution, making + the text more focused or diverse. + | Consider the following example: The AI model has to complete the + sentence “The cat is \____.” with the following token + probabilities: - Consider the following example: The AI model has to complete the - sentence “The cat is \____.” with the following token probabilities: + | playing: 0.5 + | sleeping: 0.25 + | eating: 0.15 + | driving: 0.05 + | flying: 0.05 - :: - - playing: 0.5 - sleeping: 0.25 - eating: 0.15 - driving: 0.05 - flying: 0.05 - - - **Low temperature** (e.g., 0.2): The AI model becomes more focused and deterministic, choosing tokens with the highest probability, such as "playing." - - **Medium temperature** (e.g., 1.0): The AI model maintains a balance between creativity and focus, selecting tokens based on their probabilities without significant bias, such as "playing," "sleeping," or "eating." - - **High temperature** (e.g., 2.0): The AI model becomes more adventurous, increasing the chances of selecting less likely tokens, such as "driving" and "flying." + - **Low temperature** (e.g., 0.2): The AI model becomes more focused + and deterministic, choosing tokens with the highest probability, + such as “playing.” + - **Medium temperature** (e.g., 1.0): The AI model maintains a + balance between creativity and focus, selecting tokens based on + their probabilities without significant bias, such as “playing,” + “sleeping,” or “eating.” + - **High temperature** (e.g., 2.0): The AI model becomes more + adventurous, increasing the chances of selecting less likely + tokens, such as “driving” and “flying.” - ``Top-p``, also known as nucleus sampling, is a parameter used to control the range of tokens considered by the AI model based on their @@ -231,13 +346,13 @@ There are several parameters that can control text generation quality: including those with lower probabilities, such as “driving” and “flying.” -- ``Top-k`` is an another popular sampling strategy. In comparision - with Top-P, which chooses from the smallest possible set of words - whose cumulative probability exceeds the probability P, in Top-K - sampling K most likely next words are filtered and the probability - mass is redistributed among only those K next words. In our example - with cat, if k=3, then only “playing”, “sleeping” and “eating” will - be taken into account as possible next word. +- ``Top-k`` is another popular sampling strategy. In comparison with + Top-P, which chooses from the smallest possible set of words whose + cumulative probability exceeds the probability P, in Top-K sampling K + most likely next words are filtered and the probability mass is + redistributed among only those K next words. In our example with cat, + if k=3, then only “playing”, “sleeping” and “eating” will be taken + into account as possible next word. To optimize the generation process and use memory more efficiently, the ``use_cache=True`` option is enabled. Since the output side is @@ -263,8 +378,9 @@ generated tokens without waiting until when the whole generation is finished using Streaming API, it adds a new token to the output queue and then prints them when they are ready. -Setup imports -~~~~~~~~~~~~~ +Setup imports `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -275,11 +391,12 @@ Setup imports from transformers import AutoTokenizer, TextIteratorStreamer import numpy as np -Prepare template for user prompt -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Prepare template for user prompt `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + For effective generation, model expects to have input in specific -format. The code below prepare tamplate for passing user instruction +format. The code below prepare template for passing user instruction into model with providing additional context. .. code:: ipython3 @@ -306,10 +423,11 @@ into model with providing additional context. response_key=RESPONSE_KEY, ) -Helpers for output parsing -~~~~~~~~~~~~~~~~~~~~~~~~~~ +Helpers for output parsing `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Model was retrained to finish generaton using special token ``### End`` + +Model was retrained to finish generation using special token ``### End`` the code below find its id for using it as generation stop-criteria. .. code:: ipython3 @@ -346,8 +464,9 @@ the code below find its id for using it as generation stop-criteria. except ValueError: pass -Main generation function -~~~~~~~~~~~~~~~~~~~~~~~~ +Main generation function `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + As it was discussed above, ``run_generation`` function is the entry point for starting generation. It gets provided input instruction as @@ -406,8 +525,9 @@ parameter and returns model response. start = perf_counter() return model_output, perf_text -Helpers for application -~~~~~~~~~~~~~~~~~~~~~~~ +Helpers for application `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + For making interactive user interface we will use Gradio library. The code bellow provides useful functions used for communication with UI @@ -435,7 +555,7 @@ elements. per_token_time.append(num_current_toks / current_time) if len(per_token_time) > 10 and len(per_token_time) % 4 == 0: current_bucket = per_token_time[:-10] - return f"Average generaton speed: {np.mean(current_bucket):.2f} tokens/s. Total generated tokens: {num_tokens}", num_tokens + return f"Average generation speed: {np.mean(current_bucket):.2f} tokens/s. Total generated tokens: {num_tokens}", num_tokens return current_perf_text, num_tokens def reset_textbox(instruction:str, response:str, perf:str): @@ -472,8 +592,9 @@ elements. ov_model.compile() return current_text -Run instruction-following pipeline ----------------------------------- +Run instruction-following pipeline `⇑ <#top>`__ +############################################################################################################################### + Now, we are ready to explore model capabilities. This demo provides a simple interface that allows communication with a model using text @@ -495,10 +616,7 @@ generation parameters: .. code:: ipython3 - from openvino.runtime import Core - - core = Core() - available_devices = core.available_devices + available_devices = Core().available_devices + ["AUTO"] examples = [ "Give me recipe for pizza with pineapple", @@ -559,6 +677,12 @@ generation parameters: demo.launch(enable_queue=True, share=True, height=800) +.. parsed-literal:: + + /tmp/ipykernel_1272681/896135151.py:57: GradioDeprecationWarning: The `enable_queue` parameter has been deprecated. Please use the `.queue()` method instead. + demo.launch(enable_queue=True, share=False, height=800) + + .. parsed-literal:: Running on local URL: http://127.0.0.1:7860 diff --git a/docs/notebooks/241-riffusion-text-to-music-with-output.rst b/docs/notebooks/241-riffusion-text-to-music-with-output.rst index c8af08bdae7..cae9b6e81d1 100644 --- a/docs/notebooks/241-riffusion-text-to-music-with-output.rst +++ b/docs/notebooks/241-riffusion-text-to-music-with-output.rst @@ -1,6 +1,8 @@ Text-to-Music generation using Riffusion and OpenVINO ===================================================== +.. _top: + `Riffusion `__ is a latent text-to-image diffusion model capable of generating spectrogram images given any text input. These spectrograms can be converted into @@ -36,8 +38,8 @@ About Riffusion --------------- Riffusion is based on Stable Diffusion v1.5 and fine-tuned on images of -spectrograms paired with text. Audio processing happens downstream of -the model. This model can generate an audio spectrogram for given input +spectrogram paired with text. Audio processing happens downstream of the +model. This model can generate an audio spectrogram for given input text. An audio spectrogram is a visual way to represent the frequency content @@ -50,7 +52,10 @@ represents time, and the y-axis represents frequency. The color of each pixel gives the amplitude of the audio at the frequency and time given by its row and column. -|spectrogram.png| +.. figure:: https://www.riffusion.com/_next/image?url=%2F_next%2Fstatic%2Fmedia%2Fspectrogram_label.8c8aea56.png&w=1920&q=75 + :alt: spectrogram + + spectrogram `\*image source `__ @@ -68,13 +73,23 @@ amplitudes and phases. `\*image source `__ The STFT is invertible, so the original audio can be reconstructed from -a spectrogram. This idea is a behaind approach to using Riffusion for +a spectrogram. This idea is a behind approach to using Riffusion for audio generation. -.. |spectrogram.png| image:: https://camo.githubusercontent.com/b7b275efa48f909b32dd9a5124e4eb678fde6e4f435ce7547c71366543baae62/68747470733a2f2f7777772e726966667573696f6e2e636f6d2f5f6e6578742f696d6167653f75726c3d2532465f6e6578742532467374617469632532466d656469612532467370656374726f6772616d5f6c6162656c2e38633861656135362e706e6726773d3139323026713d3735 +**Table of contents**: + +- `Prerequisites <#prerequisites>`__ +- `Stable Diffusion pipeline in Optimum Intel <#stable-diffusion-pipeline-in-optimum-intel>`__ + + - `Select inference device <#select-inference-device>`__ + +- `Prepare postprocessing for reconstruction audio from spectrogram image <#prepare-postprocessing-for-reconstruction-audio-from-spectrogram-image>`__ +- `Run Inference pipeline <#run-inference-pipeline>`__ +- `Interactive demo <#interactive-demo>`__ + +Prerequisites `⇑ <#top>`__ +############################################################################################################################### -Prerequisites -------------- .. code:: ipython3 @@ -84,9 +99,13 @@ Prerequisites .. parsed-literal:: - ERROR: pip's dependency resolver does not currently take into account all the packages that are installed. This behaviour is the source of the following dependency conflicts. - visualdl 2.5.2 requires gradio==3.11.0, but you have gradio 3.34.0 which is incompatible. + [notice] A new release of pip is available: 23.1.2 -> 23.2 + [notice] To update, run: pip install --upgrade pip + + [notice] A new release of pip is available: 23.1.2 -> 23.2 + [notice] To update, run: pip install --upgrade pip + .. code:: ipython3 @@ -97,8 +116,17 @@ Prerequisites else: !pip install -q "torchaudio==0.13.1+cpu" --find-links https://download.pytorch.org/whl/torch_stable.html -Stable Diffusion pipeline in Optimum Intel ------------------------------------------- + +.. parsed-literal:: + + + [notice] A new release of pip is available: 23.1.2 -> 23.2 + [notice] To update, run: pip install --upgrade pip + + +Stable Diffusion pipeline in Optimum Intel `⇑ <#top>`__ +############################################################################################################################### + As the riffusion model architecture is the same as Stable Diffusion, we can use it with the Stable Diffusion pipeline for text-to-image @@ -110,10 +138,10 @@ APIs. When Stable Diffusion models are exported to the OpenVINO format, they are decomposed into three components that consist of four models combined during inference into the pipeline: -* The text encoder -* The U-NET -* The VAE encoder -* The VAE decoder +- The text encoder +- The U-NET +- The VAE encoder +- The VAE decoder More details about the Stable Diffusion pipeline can be found in `stable-diffusion <225-stable-diffusion-text-to-image-with-output.html>`__ @@ -133,12 +161,44 @@ running. MODEL_ID = "riffusion/riffusion-model-v1" MODEL_DIR = Path("riffusion_pipeline") - DEVICE = "CPU" + +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + from openvino.runtime import Core + + core = Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=2, options=('CPU', 'GPU', 'AUTO'), value='AUTO') + + .. code:: ipython3 from optimum.intel.openvino import OVStableDiffusionPipeline + DEVICE = device.value + if not MODEL_DIR.exists(): pipe = OVStableDiffusionPipeline.from_pretrained(MODEL_ID, export=True, device=DEVICE, compile=False) pipe.half() @@ -149,10 +209,10 @@ running. .. parsed-literal:: - 2023-06-09 16:44:58.569326: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. - 2023-06-09 16:44:58.606428: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + 2023-07-17 16:22:33.905103: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-07-17 16:22:33.943298: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. - 2023-06-09 16:44:59.215375: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + 2023-07-17 16:22:34.567997: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT .. parsed-literal:: @@ -164,13 +224,166 @@ running. No CUDA runtime is found, using CUDA_HOME='/usr/local/cuda' comet_ml is installed but `COMET_API_KEY` is not set. - /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/diffusers/models/cross_attention.py:30: FutureWarning: Importing from cross_attention is deprecated. Please import from diffusers.models.attention_processor instead. - deprecate( - The config attributes {'safety_checker': ['stable_diffusion', 'StableDiffusionSafetyChecker']} were passed to OVStableDiffusionPipeline, but are not expected and will be ignored. Please verify your model_index.json configuration file. -Prepare postprocessing for reconstruction audio from spectrogram image ----------------------------------------------------------------------- + +.. parsed-literal:: + + Downloading (…)ain/model_index.json: 0%| | 0.00/541 [00:00= 64: + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/diffusers/models/unet_2d_condition.py:977: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + if not return_dict: + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/torch/onnx/_internal/jit_utils.py:258: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + _C._jit_pass_onnx_node_shape_type_inference(node, params_dict, opset_version) + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + _C._jit_pass_onnx_graph_shape_type_inference( + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + _C._jit_pass_onnx_graph_shape_type_inference( + Saving external data to one file... + Using framework PyTorch: 1.13.1+cpu + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/torch/onnx/_internal/jit_utils.py:258: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) + _C._jit_pass_onnx_node_shape_type_inference(node, params_dict, opset_version) + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) + _C._jit_pass_onnx_graph_shape_type_inference( + /home/ea/work/notebooks_convert/notebooks_conv_env/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) + _C._jit_pass_onnx_graph_shape_type_inference( + Using framework PyTorch: 1.13.1+cpu + + +Prepare postprocessing for reconstruction audio from spectrogram image. `⇑ <#top>`__ +############################################################################################################################### The riffusion model generates an audio spectrogram image, which can be used to reconstruct audio. However, the spectrogram images from the @@ -192,7 +405,7 @@ scale `__, which is a perceptual scale of pitches judged by listeners to be equal in distance from one another. -The code below defines the process of reconstruction of a wav audio clip +The code below defines the process of reconstruction of a WAV audio clip from a spectrogram image using Griffin-Lim Algorithm. .. code:: ipython3 @@ -336,8 +549,9 @@ from a spectrogram image using Griffin-Lim Algorithm. return waveform -Run Inference pipeline ----------------------- +Run Inference pipeline `⇑ <#top>`__ +############################################################################################################################### + The diagram below briefly describes the workflow of our pipeline @@ -349,7 +563,7 @@ The diagram below briefly describes the workflow of our pipeline As you can see, it is very similar to Stable Diffusion Text-to-Image generation with an additional post-processing step that transforms generated spectrogram into an audio signal. Firstly, -OVStableDiffusionPipeline accepts input text prompt, which will be +``OVStableDiffusionPipeline`` accepts input text prompt, which will be tokenized and transformed to embeddings space using Frozen CLIP text encoder and generates initial latent spectrogram representation using a random generator, then U-Net iteratively *denoises* the random latent @@ -390,9 +604,9 @@ reconstructed audio. .. parsed-literal:: - Compiling the encoder and creating the inference request ... - Compiling the encoder and creating the inference request ... - Compiling the encoder and creating the inference request ... + Compiling the text_encoder... + Compiling the vae_decoder... + Compiling the unet... Now, we can test our generation. Function generate accepts text input @@ -423,7 +637,7 @@ without the other. More explanation of how it works can be found in this -.. image:: 241-riffusion-text-to-music-with-output_files/241-riffusion-text-to-music-with-output_13_0.png +.. image:: 241-riffusion-text-to-music-with-output_files/241-riffusion-text-to-music-with-output_15_0.png @@ -439,22 +653,23 @@ without the other. More explanation of how it works can be found in this -Interactive demo ----------------- +Interactive demo `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 import gradio as gr from openvino.runtime import Core - available_devices = Core().available_devices + available_devices = Core().available_devices + ["AUTO"] examples = [ "acoustic folk violin jam", @@ -520,7 +735,15 @@ Interactive demo .. parsed-literal:: - Running on local URL: http://127.0.0.1:7860 + /tmp/ipykernel_1282292/2438576232.py:56: GradioDeprecationWarning: The `style` method is deprecated. Please set these arguments in the constructor instead. + spectrogram_output.style(height=256) + /tmp/ipykernel_1282292/2438576232.py:63: GradioDeprecationWarning: The `enable_queue` parameter has been deprecated. Please use the `.queue()` method instead. + demo.launch(enable_queue=True, height=800) + + +.. parsed-literal:: + + Running on local URL: http://127.0.0.1:7861 To create a public link, set `share=True` in `launch()`. @@ -528,11 +751,5 @@ Interactive demo .. raw:: html -
- - - -.. parsed-literal:: - - 0%| | 0/21 [00:00 diff --git a/docs/notebooks/241-riffusion-text-to-music-with-output_files/241-riffusion-text-to-music-with-output_13_0.png b/docs/notebooks/241-riffusion-text-to-music-with-output_files/241-riffusion-text-to-music-with-output_13_0.png deleted file mode 100644 index a305da89434..00000000000 --- a/docs/notebooks/241-riffusion-text-to-music-with-output_files/241-riffusion-text-to-music-with-output_13_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:98185d77549537a8562a2c13222c78932cf0458be457e326f6937efb3ec6d3d9 -size 514408 diff --git a/docs/notebooks/241-riffusion-text-to-music-with-output_files/241-riffusion-text-to-music-with-output_15_0.jpg b/docs/notebooks/241-riffusion-text-to-music-with-output_files/241-riffusion-text-to-music-with-output_15_0.jpg new file mode 100644 index 00000000000..605e464bf82 --- /dev/null +++ b/docs/notebooks/241-riffusion-text-to-music-with-output_files/241-riffusion-text-to-music-with-output_15_0.jpg @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:eab8042875afdad8eb30bf5edd079896eb72338a3c65d682dc7d6aac489bc354 +size 55481 diff --git a/docs/notebooks/241-riffusion-text-to-music-with-output_files/241-riffusion-text-to-music-with-output_15_0.png b/docs/notebooks/241-riffusion-text-to-music-with-output_files/241-riffusion-text-to-music-with-output_15_0.png new file mode 100644 index 00000000000..39b7e78b64a --- /dev/null +++ b/docs/notebooks/241-riffusion-text-to-music-with-output_files/241-riffusion-text-to-music-with-output_15_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:41d293c15a42032bfa5b9f22f808f31b127276aae59afdd1fd36320a11e747ba +size 493898 diff --git a/docs/notebooks/241-riffusion-text-to-music-with-output_files/index.html b/docs/notebooks/241-riffusion-text-to-music-with-output_files/index.html index 8382ba2b8e4..495e93bde31 100644 --- a/docs/notebooks/241-riffusion-text-to-music-with-output_files/index.html +++ b/docs/notebooks/241-riffusion-text-to-music-with-output_files/index.html @@ -1,7 +1,8 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/241-riffusion-text-to-music-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/241-riffusion-text-to-music-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/241-riffusion-text-to-music-with-output_files/


../
-241-riffusion-text-to-music-with-output_13_0.png   12-Jul-2023 00:11              514408
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/241-riffusion-text-to-music-with-output_files/


../
+241-riffusion-text-to-music-with-output_15_0.jpg   16-Aug-2023 01:31               55481
+241-riffusion-text-to-music-with-output_15_0.png   16-Aug-2023 01:31              493898
 

diff --git a/docs/notebooks/242-freevc-voice-conversion-with-output.rst b/docs/notebooks/242-freevc-voice-conversion-with-output.rst index 1e556ac9790..30c08385bd5 100644 --- a/docs/notebooks/242-freevc-voice-conversion-with-output.rst +++ b/docs/notebooks/242-freevc-voice-conversion-with-output.rst @@ -1,5 +1,7 @@ -High-Quality Text-Free One-Shot Voice Conversion with FeeVC and OpenVINO™ -========================================================================= +High-Quality Text-Free One-Shot Voice Conversion with FreeVC and OpenVINO™ +========================================================================== + +.. _top: `FreeVC `__ allows alter the voice of a source speaker to a target style, while keeping the linguistic content @@ -17,27 +19,35 @@ flow. Detailed information is available in this Inference -`image_source `__ +`\**image_source\* `__ FreeVC suggests only command line interface to use and only with CUDA. In this notebook it shows how to use FreeVC in Python and without CUDA devices. It consists of the following steps: -- Download and prepare models. -- Inference. -- Convert models to OpenVINO Intermediate Representation. -- Inference using only OpenVINO's IR models. +- Download and prepare models. +- Inference. +- Convert models to OpenVINO Intermediate Representation. +- Inference using only OpenVINO’s IR models. +**Table of contents**: -Pre-requisites --------------- +- `Prerequisites <#prerequisites>`__ +- `Imports and settings <#imports-and-settings>`__ +- `Convert Modes to OpenVINO Intermediate Representation <#convert-modes-to-openvino-intermediate-representation>`__ -This steps can be done manually or will be performed automatically -during the execution of the notebook, but in minimum necessary scope. 1. -Clone this repo: git clone https://github.com/OlaWod/FreeVC.git. 2. -Download + - `Convert Prior Encoder. <#convert-prior-encoder>`__ + - `Convert SpeakerEncoder <#convert-speakerencoder>`__ + - `Convert Decoder <#convert-decoder>`__ + +Prerequisites `⇑ <#top>`__ +############################################################################################################################### + +This steps can be done manually or will be performed automatically during the execution of the notebook, but in +minimum necessary scope. 1. Clone this repo: git clone +https://github.com/OlaWod/FreeVC.git. 2. Download `WavLM-Large `__ -and put it under directory ‘FreeVC/wavlm/’. 3. You can download the +and put it under directory ``FreeVC/wavlm/``. 3. You can download the `VCTK `__ dataset. For this example we download only two of them from `Hugging Face FreeVC example `__. 4. @@ -54,7 +64,15 @@ Install extra requirements !pip install -q "webrtcvad==2.0.10" !pip install -q gradio -Check if FreeVC is installed and its path to sys.path + +.. parsed-literal:: + + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + + +Check if FreeVC is installed and append its path to ``sys.path`` .. code:: ipython3 @@ -73,11 +91,11 @@ Check if FreeVC is installed and its path to sys.path Cloning into 'FreeVC'... remote: Enumerating objects: 131, done. - remote: Counting objects: 100% (47/47), done. - remote: Compressing objects: 100% (30/30), done. - remote: Total 131 (delta 29), reused 19 (delta 17), pack-reused 84 - Receiving objects: 100% (131/131), 15.28 MiB | 3.83 MiB/s, done. - Resolving deltas: 100% (39/39), done. + remote: Counting objects: 100% (61/61), done. + remote: Compressing objects: 100% (40/40), done. + remote: Total 131 (delta 36), reused 21 (delta 21), pack-reused 70 + Receiving objects: 100% (131/131), 15.28 MiB | 4.14 MiB/s, done. + Resolving deltas: 100% (43/43), done. .. code:: ipython3 @@ -146,8 +164,9 @@ Check if FreeVC is installed and its path to sys.path p226_002.wav: 0%| | 0.00/135k [00:00`__ +############################################################################################################################### + .. code:: ipython3 @@ -173,7 +192,7 @@ Imports and settings logger = logging.getLogger() logger.setLevel(logging.CRITICAL) -Redefine function ``get_model`` from ``utils`` to exclude cuda +Redefine function ``get_model`` from ``utils`` to exclude CUDA .. code:: ipython3 @@ -206,7 +225,7 @@ Models initialization .. parsed-literal:: - Loaded the voice encoder model on cpu in 0.00 seconds. + Loaded the voice encoder model on cpu in 0.01 seconds. Reading dataset settings @@ -246,34 +265,36 @@ Inference .. parsed-literal:: - 2it [00:01, 1.34it/s] + 2it [00:01, 1.27it/s] Result audio files should be available in ‘outputs/freevc’ -Convert Modes to OpenVINO Intermediate Representation -===================================================== +Convert Modes to OpenVINO Intermediate Representation `⇑ <#top>`__ +#################################################################### -Convert each model to ONNX format and then call the OpenVINO Model -Optimizer Python API to convert the ONNX model to OpenVINO IR, with FP16 -precision. ``mo.convert_model`` function accept path to a model and -returns OpenVINO Model class instance which represents this model. -Obtained model is ready to use and loading on device using -``compile_model`` or can be saved on disk using ``serialize`` function. -``read_model`` method loads a saved model from disk. See the `Model -Optimizer Developer -Guide `__ -for more information about Model Optimizer. ### Convert Prior Encoder. -First we convert WavLM model, as a part of Convert Prior Encoder, to the -ONNX format, then to OpenVINO’s IR format. And we keep original name of -model in code: ``cmodel``. +Convert each model to ONNX format and then use the model conversion +Python API to convert the ONNX model to OpenVINO IR, with FP16 +precision. The ``mo.convert_model`` function accepts the path to a model +and returns the OpenVINO Model class instance which represents this +model. The obtained model is ready to use and to be loaded on a device +using ``compile_model`` or can be saved on a disk using the +``serialize`` function. The ``read_model`` method loads a saved model +from a disk. For more information about model conversion, see this +`page `__. + +Convert Prior Encoder. `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + +First we convert WavLM model, as a part of Convert Prior Encoder, to the ONNX format, then to OpenVINO’s IR +format. We keep the original name of the model in code: ``cmodel``. .. code:: ipython3 # define forward as extract_features for compatibility cmodel.forward = cmodel.extract_features -Convert cmodel to ONNX. +Convert ``cmodel`` to ONNX. .. code:: ipython3 @@ -308,19 +329,19 @@ Convert cmodel to ONNX. .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/242-freevc-voice-conversion/FreeVC/wavlm/WavLM.py:352: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/242-freevc-voice-conversion/FreeVC/wavlm/WavLM.py:352: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if mask: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/242-freevc-voice-conversion/FreeVC/wavlm/modules.py:495: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/242-freevc-voice-conversion/FreeVC/wavlm/modules.py:495: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! assert embed_dim == self.embed_dim - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/242-freevc-voice-conversion/FreeVC/wavlm/modules.py:496: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/242-freevc-voice-conversion/FreeVC/wavlm/modules.py:496: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! assert list(query.size()) == [tgt_len, bsz, embed_dim] - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/242-freevc-voice-conversion/FreeVC/wavlm/modules.py:500: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/242-freevc-voice-conversion/FreeVC/wavlm/modules.py:500: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! assert key_bsz == bsz - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/242-freevc-voice-conversion/FreeVC/wavlm/modules.py:502: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/242-freevc-voice-conversion/FreeVC/wavlm/modules.py:502: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! assert src_len, bsz == value.shape[:2] - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/242-freevc-voice-conversion/FreeVC/wavlm/WavLM.py:372: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/242-freevc-voice-conversion/FreeVC/wavlm/WavLM.py:372: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! feature = res["features"] if ret_conv else res["x"] - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/242-freevc-voice-conversion/FreeVC/wavlm/WavLM.py:373: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/242-freevc-voice-conversion/FreeVC/wavlm/WavLM.py:373: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! if ret_layer_results: @@ -335,11 +356,38 @@ Converting to OpenVINO’s IR format. serialize(ir_cmodel, str(ir_cmodel_path)) else: ir_cmodel = core.read_model(ir_cmodel_path) - - compiled_cmodel = core.compile_model(ir_cmodel, 'CPU') -Convert SpeakerEncoder -~~~~~~~~~~~~~~~~~~~~~~ +Select device from dropdown list for running inference using OpenVINO + +.. code:: ipython3 + + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + + compiled_cmodel = core.compile_model(ir_cmodel, device.value) + +Convert ``SpeakerEncoder`` `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ Converting to ONNX format. @@ -372,13 +420,13 @@ Converting to ONNX format. .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/symbolic_opset9.py:4315: UserWarning: Exporting a model to ONNX with a batch_size other than 1, with a variable length with LSTM can cause an error when running the ONNX model with a different batch size. Make sure to save the model with a batch size of 1, or define the initial states (h0/c0) as inputs of the model. + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/symbolic_opset9.py:4315: UserWarning: Exporting a model to ONNX with a batch_size other than 1, with a variable length with LSTM can cause an error when running the ONNX model with a different batch size. Make sure to save the model with a batch size of 1, or define the initial states (h0/c0) as inputs of the model. warnings.warn( - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/_internal/jit_utils.py:258: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/_internal/jit_utils.py:258: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) _C._jit_pass_onnx_node_shape_type_inference(node, params_dict, opset_version) - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) _C._jit_pass_onnx_graph_shape_type_inference( - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) _C._jit_pass_onnx_graph_shape_type_inference( @@ -501,17 +549,33 @@ based on ``speaker_encoder.voice_encoder.SpeakerEncoder`` class methods return embed, partial_embeds, wav_slices return embed +Select device from dropdown list for running inference using OpenVINO + +.. code:: ipython3 + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + Then compile model. .. code:: ipython3 - compiled_smodel = core.compile_model(ir_smodel, 'CPU') + compiled_smodel = core.compile_model(ir_smodel, device.value) -Convert Decoder -~~~~~~~~~~~~~~~ +Convert Decoder `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -In the same way export SynthesizerTrn model, that implements decoder -function, to ONNX format and convert it to OpenVino IR format. + +In the same way export ``SynthesizerTrn`` model, that implements decoder +function, to ONNX format and convert it to OpenVINO IR format. .. code:: ipython3 @@ -546,8 +610,12 @@ function, to ONNX format and convert it to OpenVino IR format. serialize(ir_net_g_model, str(ir_net_g_path)) else: ir_net_g_model = core.read_model(ir_net_g_path) - - compiled_ir_net_g_model = core.compile_model(ir_net_g_model, 'CPU') + +Select device from dropdown list for running inference using OpenVINO + +.. code:: ipython3 + + compiled_ir_net_g_model = core.compile_model(ir_net_g_model, device.value) Define function for synthesizing. @@ -598,7 +666,7 @@ And now we can check inference using only IR models. .. parsed-literal:: - 2it [00:02, 1.48s/it] + 2it [00:02, 1.45s/it] Result audio files should be available in ‘outputs/freevc’ and you can @@ -659,7 +727,7 @@ Result audio: @@ -705,15 +773,15 @@ inference. Use rate corresponding to the value of .. parsed-literal:: - /tmp/ipykernel_3466225/3932271335.py:4: GradioDeprecationWarning: Usage of gradio.inputs is deprecated, and will not be supported in the future, please import your component from gradio.components + /tmp/ipykernel_2082705/3932271335.py:4: GradioDeprecationWarning: Usage of gradio.inputs is deprecated, and will not be supported in the future, please import your component from gradio.components audio1 = gr.inputs.Audio(label="Source Audio", type='filepath') - /tmp/ipykernel_3466225/3932271335.py:4: GradioDeprecationWarning: `optional` parameter is deprecated, and it has no effect + /tmp/ipykernel_2082705/3932271335.py:4: GradioDeprecationWarning: `optional` parameter is deprecated, and it has no effect audio1 = gr.inputs.Audio(label="Source Audio", type='filepath') - /tmp/ipykernel_3466225/3932271335.py:5: GradioDeprecationWarning: Usage of gradio.inputs is deprecated, and will not be supported in the future, please import your component from gradio.components + /tmp/ipykernel_2082705/3932271335.py:5: GradioDeprecationWarning: Usage of gradio.inputs is deprecated, and will not be supported in the future, please import your component from gradio.components audio2 = gr.inputs.Audio(label="Reference Audio", type='filepath') - /tmp/ipykernel_3466225/3932271335.py:5: GradioDeprecationWarning: `optional` parameter is deprecated, and it has no effect + /tmp/ipykernel_2082705/3932271335.py:5: GradioDeprecationWarning: `optional` parameter is deprecated, and it has no effect audio2 = gr.inputs.Audio(label="Reference Audio", type='filepath') - /tmp/ipykernel_3466225/3932271335.py:6: GradioDeprecationWarning: Usage of gradio.outputs is deprecated, and will not be supported in the future, please import your components from gradio.components + /tmp/ipykernel_2082705/3932271335.py:6: GradioDeprecationWarning: Usage of gradio.outputs is deprecated, and will not be supported in the future, please import your components from gradio.components outputs = gr.outputs.Audio(label="Output Audio", type='filepath') diff --git a/docs/notebooks/243-tflite-selfie-segmentation-with-output.rst b/docs/notebooks/243-tflite-selfie-segmentation-with-output.rst index f5d0cdb5645..9e2ba45e11e 100644 --- a/docs/notebooks/243-tflite-selfie-segmentation-with-output.rst +++ b/docs/notebooks/243-tflite-selfie-segmentation-with-output.rst @@ -1,6 +1,8 @@ Selfie Segmentation using TFLite and OpenVINO ============================================= +.. _top: + The Selfie segmentation pipeline allows developers to easily separate the background from users within a scene and focus on what matters. Adding cool effects to selfies or inserting your users into interesting @@ -12,7 +14,7 @@ In this tutorial, we consider how to implement selfie segmentation using OpenVINO. We will use `Multiclass Selfie-segmentation model `__ provided as part of `Google -Mediapipe `__ solution. +MediaPipe `__ solution. The Multiclass Selfie-segmentation model is a multiclass semantic segmentation model and classifies each pixel as background, hair, body, @@ -34,16 +36,43 @@ The tutorial consists of following steps: 2. Run inference on the image. 3. Run interactive background blurring demo on video. -Prerequisites -------------- +**Table of contents**: + +- `Prerequisites <#prerequisites>`__ + + - `Install required dependencies <#install-required-dependencies>`__ + - `Download pre-trained model and test image <#download-pre-trained-model-and-test-image>`__ + +- `Convert Tensorflow Lite model to OpenVINO IR format <#convert-tensorflow-lite-model-to-openvino-ir-format>`__ +- `Run OpenVINO model inference on image <#run-openvino-model-inference-on-image>`__ + + - `Load model <#load-model>`__ + - `Prepare input image <#prepare-input-image>`__ + - `Run model inference <#run-model-inference>`__ + - `Postprocess and visualize inference results <#postprocess-and-visualize-inference-results>`__ + +- `Interactive background blurring demo on video <#interactive-background-blurring-demo-on-video>`__ + + - `Run Live Background Blurring <#run-live-background-blurring>`__ + +Prerequisites `⇑ <#top>`__ +############################################################################################################################### + + +Install required dependencies `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Install required dependencies -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ .. code:: ipython3 !pip install -q "openvino-dev>=2023.0.0" "matplotlib" "opencv-python" + +.. parsed-literal:: + + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + + .. code:: ipython3 import urllib.request @@ -52,8 +81,9 @@ Install required dependencies filename='notebook_utils.py' ); -Download pretrained model and test image -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Download pretrained model and test image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -76,12 +106,13 @@ Download pretrained model and test image .. parsed-literal:: - PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/243-tflite-selfie-segmentation/selfie_multiclass_256x256.tflite') + PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/243-tflite-selfie-segmentation/selfie_multiclass_256x256.tflite') -Convert Tensorflow Lite model to OpenVINO IR format ---------------------------------------------------- +Convert Tensorflow Lite model to OpenVINO IR format `⇑ <#top>`__ +############################################################################################################################### + Starting from the 2023.0.0 release, OpenVINO supports TFLite model conversion. However TFLite model format can be directly passed in @@ -92,18 +123,19 @@ tutorial with `basic OpenVINO API capabilities <002-openvino-api-with-output.html>`__), it is recommended to convert model to OpenVINO Intermediate Representation format to apply additional optimizations (e.g. weights compression to -FP16 format). To convert the TFLite model to OpenVINO IR, OpenVINO Model -Optimizer Python API can be used. ``mo.convert_model`` function accepts -a path to the TFLite model and returns the OpenVINO Model class instance -which represents this model. The obtained model is ready to use and -loading on the device using ``compile_model`` or can be saved on disk -using the ``serialize`` function reducing loading time for the next -running. Optionally, we can apply compression to FP16 model weights -using ``compress_to_fp16=True`` option and integrate preprocessing using -this approach. See the `Model Optimizer Developer -Guide `__ -for more information about Model Optimizer and TensorFlow Lite `models -suport `__. +FP16 format). To convert the TFLite model to OpenVINO IR, model +conversion Python API can be used. The ``mo.convert_model`` function +accepts a path to the TFLite model and returns the OpenVINO Model class +instance which represents this model. The obtained model is ready to use +and to be loaded on the device using ``compile_model`` or can be saved +on a disk using the ``serialize`` function reducing loading time for the +next running. Optionally, we can apply compression to the FP16 model +weights, using the ``compress_to_fp16=True`` option and integrate +preprocessing, using this approach. For more information about model +conversion, see this +`page `__. +For TensorFlow Lite, refer to the `models +support `__. .. code:: ipython3 @@ -152,21 +184,23 @@ division on 255. Model output is a floating point tensor with the similar format and -shape, except number of channels - 6 that epresents number of supported +shape, except number of channels - 6 that represents number of supported segmentation classes: background, hair, body skin, face skin, clothes, and others. Each value in the output tensor represents of probability that the pixel belongs to the specified class. We can use the ``argmax`` operation to get the label with the highest probability for each pixel. -Run OpenVINO model inference on image -------------------------------------- +Run OpenVINO model inference on image `⇑ <#top>`__ +############################################################################################################################### + Let’s see the model in action. For running the inference model with OpenVINO we should load the model on the device first. Please use the next dropdown list for the selection inference device. -Load model -~~~~~~~~~~ +Load model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -194,8 +228,9 @@ Load model compiled_model = core.compile_model(ov_model, device.value) -Prepare input image -~~~~~~~~~~~~~~~~~~~ +Prepare input image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The model accepts an image with size 256x256, we need to resize our input image to fit it in the model input tensor. Usually, segmentation @@ -249,15 +284,17 @@ Additionally, the input image is represented as an RGB image in UINT8 # Convert input data from uint8 [0, 255] to float32 [0, 1] range and add batch dimension normalized_img = np.expand_dims(padded_img.astype(np.float32) / 255, 0) -Run model inference -~~~~~~~~~~~~~~~~~~~ +Run model inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 out = compiled_model(normalized_img)[0] -Posprocess and visualize inference results -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Postprocess and visualize inference results `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The model predicts segmentation probabilities mask with the size 256 x 256, we need to apply postprocessing to get labels with the highest @@ -359,11 +396,12 @@ Visualize obtained result -.. image:: 243-tflite-selfie-segmentation-with-output_files/243-tflite-selfie-segmentation-with-output_24_0.png +.. image:: 243-tflite-selfie-segmentation-with-output_files/243-tflite-selfie-segmentation-with-output_25_0.png -Interactive background blurring demo on video ---------------------------------------------- +Interactive background blurring demo on video `⇑ <#top>`__ +############################################################################################################################### + The following code runs model inference on a video: @@ -485,8 +523,9 @@ The following code runs model inference on a video: if use_popup: cv2.destroyAllWindows() -Run Live Background Blurring -~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Run Live Background Blurring `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Use a webcam as the video input. By default, the primary webcam is set with \ ``source=0``. If you have multiple webcams, each one will be @@ -534,7 +573,7 @@ Run: -.. image:: 243-tflite-selfie-segmentation-with-output_files/243-tflite-selfie-segmentation-with-output_32_0.png +.. image:: 243-tflite-selfie-segmentation-with-output_files/243-tflite-selfie-segmentation-with-output_33_0.png .. parsed-literal:: diff --git a/docs/notebooks/243-tflite-selfie-segmentation-with-output_files/243-tflite-selfie-segmentation-with-output_24_0.png b/docs/notebooks/243-tflite-selfie-segmentation-with-output_files/243-tflite-selfie-segmentation-with-output_25_0.png similarity index 100% rename from docs/notebooks/243-tflite-selfie-segmentation-with-output_files/243-tflite-selfie-segmentation-with-output_24_0.png rename to docs/notebooks/243-tflite-selfie-segmentation-with-output_files/243-tflite-selfie-segmentation-with-output_25_0.png diff --git a/docs/notebooks/243-tflite-selfie-segmentation-with-output_files/243-tflite-selfie-segmentation-with-output_32_0.png b/docs/notebooks/243-tflite-selfie-segmentation-with-output_files/243-tflite-selfie-segmentation-with-output_32_0.png deleted file mode 100644 index 1d656e9cad6..00000000000 --- a/docs/notebooks/243-tflite-selfie-segmentation-with-output_files/243-tflite-selfie-segmentation-with-output_32_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:3367b6793598ce577cecd266da289d04d0b93163b4c0c05fcfb23fdf16eb5eaa -size 14406 diff --git a/docs/notebooks/243-tflite-selfie-segmentation-with-output_files/243-tflite-selfie-segmentation-with-output_33_0.png b/docs/notebooks/243-tflite-selfie-segmentation-with-output_files/243-tflite-selfie-segmentation-with-output_33_0.png new file mode 100644 index 00000000000..a35b90fd788 --- /dev/null +++ b/docs/notebooks/243-tflite-selfie-segmentation-with-output_files/243-tflite-selfie-segmentation-with-output_33_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:3b6e654bff9d785a50a4a40d4ae7b06168a9439fda6d2e77ed927030cb7a8f68 +size 14356 diff --git a/docs/notebooks/243-tflite-selfie-segmentation-with-output_files/index.html b/docs/notebooks/243-tflite-selfie-segmentation-with-output_files/index.html index 8eadf2110c9..6bb3be94cee 100644 --- a/docs/notebooks/243-tflite-selfie-segmentation-with-output_files/index.html +++ b/docs/notebooks/243-tflite-selfie-segmentation-with-output_files/index.html @@ -1,8 +1,8 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/243-tflite-selfie-segmentation-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/243-tflite-selfie-segmentation-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/243-tflite-selfie-segmentation-with-output_files/


../
-243-tflite-selfie-segmentation-with-output_24_0..> 12-Jul-2023 00:11              512588
-243-tflite-selfie-segmentation-with-output_32_0..> 12-Jul-2023 00:11               14406
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/243-tflite-selfie-segmentation-with-output_files/


../
+243-tflite-selfie-segmentation-with-output_25_0..> 16-Aug-2023 01:31              512588
+243-tflite-selfie-segmentation-with-output_33_0..> 16-Aug-2023 01:31               14356
 

diff --git a/docs/notebooks/244-named-entity-recognition-with-output.rst b/docs/notebooks/244-named-entity-recognition-with-output.rst new file mode 100644 index 00000000000..40dcb1455d7 --- /dev/null +++ b/docs/notebooks/244-named-entity-recognition-with-output.rst @@ -0,0 +1,508 @@ +Named entity recognition with OpenVINO™ +======================================= + +.. _top: + +The Named Entity Recognition(NER) is a natural language processing +method that involves the detecting of key information in the +unstructured text and categorizing it into pre-defined categories. These +categories or named entities refer to the key subjects of text, such as +names, locations, companies and etc. + +NER is a good method for the situations when a high-level overview of a +large amount of text is needed. NER can be helpful with such task as +analyzing key information in unstructured text or automates the +information extraction of large amounts of data. + +This tutorial shows how to perform named entity recognition using +OpenVINO. We will use the pre-trained model +`elastic/distilbert-base-cased-finetuned-conll03-english `__. +It is DistilBERT based model, trained on +`conll03 english dataset `__. +The model can recognize four named entities in text: persons, locations, +organizations and names of miscellaneous entities that do not belong to +the previous three groups. The model is sensitive to capital letters. + +To simplify the user experience, the `Hugging Face +Optimum `__ library is used to +convert the model to OpenVINO™ IR format and quantize it. + +**Table of contents**: + +- `Prerequisites <#prerequisites>`__ +- `Download the NER model <#download-the-ner-model>`__ +- `Quantize the model, using Hugging Face Optimum API <#quantize-the-model-using-hugging-face-optimum-api>`__ +- `Prepare demo for Named Entity Recognition OpenVINO Runtime <#prepare-demo-for-named-entity-recognition-openvino-runtime>`__ +- `Compare the Original and Quantized Models <#compare-the-original-and-quantized-models>`__ + + - `Compare performance <#compare-performance>`__ + - `Compare size of the models <#compare-size-of-the-models>`__ + +Prerequisites `⇑ <#top>`__ +############################################################################################################################### + + +.. code:: ipython3 + + !pip install -q "diffusers>=0.17.1" "openvino-dev>=2023.0.0" "nncf>=2.5.0" "gradio" "onnx>=1.11.0" "onnxruntime>=1.14.0" "optimum-intel>=1.9.1" "transformers>=4.31.0" + + +.. parsed-literal:: + + ERROR: pip's dependency resolver does not currently take into account all the packages that are installed. This behaviour is the source of the following dependency conflicts. + audiocraft 0.0.2a2 requires xformers, which is not installed. + audiocraft 0.0.2a2 requires torch>=2.0.0, but you have torch 1.13.1+cpu which is incompatible. + audiocraft 0.0.2a2 requires torchaudio>=2.0.0, but you have torchaudio 0.13.1+cpu which is incompatible. + deepfloyd-if 1.0.2rc0 requires accelerate~=0.15.0, but you have accelerate 0.22.0.dev0 which is incompatible. + deepfloyd-if 1.0.2rc0 requires diffusers~=0.16.0, but you have diffusers 0.18.2 which is incompatible. + deepfloyd-if 1.0.2rc0 requires transformers~=4.25.1, but you have transformers 4.30.2 which is incompatible. + paddleclas 2.5.1 requires faiss-cpu==1.7.1.post2, but you have faiss-cpu 1.7.4 which is incompatible. + paddleclas 2.5.1 requires gast==0.3.3, but you have gast 0.4.0 which is incompatible. + ppgan 2.1.0 requires librosa==0.8.1, but you have librosa 0.9.2 which is incompatible. + ppgan 2.1.0 requires opencv-python<=4.6.0.66, but you have opencv-python 4.7.0.72 which is incompatible. + pytorch-lightning 1.6.5 requires protobuf<=3.20.1, but you have protobuf 3.20.3 which is incompatible. + spacy 3.5.2 requires pydantic!=1.8,!=1.8.1,<1.11.0,>=1.7.4, but you have pydantic 2.0.3 which is incompatible. + thinc 8.1.10 requires pydantic!=1.8,!=1.8.1,<1.11.0,>=1.7.4, but you have pydantic 2.0.3 which is incompatible. + visualdl 2.5.2 requires gradio==3.11.0, but you have gradio 3.36.1 which is incompatible. + + [notice] A new release of pip is available: 23.1.2 -> 23.2 + [notice] To update, run: pip install --upgrade pip + + +Download the NER model `⇑ <#top>`__ +############################################################################################################################### + + +We load the +`distilbert-base-cased-finetuned-conll03-english `__ +model from the `Hugging Face Hub `__ with +`Hugging Face Transformers +library `__. + +Model class initialization starts with calling ``from_pretrained`` +method. To easily save the model, you can use the ``save_pretrained()`` +method. + +.. code:: ipython3 + + from transformers import AutoTokenizer, AutoModelForTokenClassification + + model_id = "elastic/distilbert-base-cased-finetuned-conll03-english" + model = AutoModelForTokenClassification.from_pretrained(model_id) + + original_ner_model_dir = 'original_ner_model' + model.save_pretrained(original_ner_model_dir) + + tokenizer = AutoTokenizer.from_pretrained(model_id) + + + +.. parsed-literal:: + + Downloading (…)lve/main/config.json: 0%| | 0.00/954 [00:00`__ +############################################################################################################################### + + +Post-training static quantization introduces an additional calibration +step where data is fed through the network in order to compute the +activations quantization parameters. For quantization it will be used +`Hugging Face Optimum Intel +API `__. + +To handle the NNCF quantization process we use class +`OVQuantizer `__. +The quantization with Hugging Face Optimum Intel API contains the next +steps: \* Model class initialization starts with calling +``from_pretrained()`` method. \* Next we create calibration dataset with +``get_calibration_dataset()`` to use for the post-training static +quantization calibration step. \* After we quantize a model and save the +resulting model in the OpenVINO IR format to save_directory with +``quantize()`` method. \* Then we load the quantized model. The Optimum +Inference models are API compatible with Hugging Face Transformers +models and we can just replace ``AutoModelForXxx`` class with the +corresponding ``OVModelForXxx`` class. So we use +``OVModelForTokenClassification`` to load the model. + +.. code:: ipython3 + + from functools import partial + from optimum.intel import OVQuantizer + + from optimum.intel import OVModelForTokenClassification + + def preprocess_fn(data, tokenizer): + examples = [] + for data_chunk in data["tokens"]: + examples.append(' '.join(data_chunk)) + + return tokenizer( + examples, padding=True, truncation=True, max_length=128 + ) + + quantizer = OVQuantizer.from_pretrained(model) + calibration_dataset = quantizer.get_calibration_dataset( + "conll2003", + preprocess_function=partial(preprocess_fn, tokenizer=tokenizer), + num_samples=100, + dataset_split="train", + preprocess_batch=True, + ) + + # The directory where the quantized model will be saved + quantized_ner_model_dir = "quantized_ner_model" + + # Apply static quantization and save the resulting model in the OpenVINO IR format + quantizer.quantize(calibration_dataset=calibration_dataset, save_directory=quantized_ner_model_dir) + + # Load the quantized model + optimized_model = OVModelForTokenClassification.from_pretrained(quantized_ner_model_dir) + + +.. parsed-literal:: + + 2023-07-17 14:40:49.402855: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-07-17 14:40:49.442756: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. + 2023-07-17 14:40:50.031065: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + + +.. parsed-literal:: + + INFO:nncf:NNCF initialized successfully. Supported frameworks detected: torch, tensorflow, onnx, openvino + + +.. parsed-literal:: + + No CUDA runtime is found, using CUDA_HOME='/usr/local/cuda' + comet_ml is installed but `COMET_API_KEY` is not set. + + + +.. parsed-literal:: + + Downloading builder script: 0%| | 0.00/9.57k [00:00`__ +############################################################################################################################### + + +As the Optimum Inference models are API compatible with Hugging Face +Transformers models, we can just use ``pipleine()`` from `Hugging Face +Transformers API `__ for +inference. + +.. code:: ipython3 + + from transformers import pipeline + + ner_pipeline_optimized = pipeline("token-classification", model=optimized_model, tokenizer=tokenizer) + +Now, you can try NER model on own text. Put your sentence to input text +box, click Submit button, the model label the recognized entities in the +text. + +.. code:: ipython3 + + import gradio as gr + + examples = [ + "My name is Wolfgang and I live in Berlin.", + ] + + def run_ner(text): + output = ner_pipeline_optimized(text) + return {"text": text, "entities": output} + + demo = gr.Interface(run_ner, + gr.Textbox(placeholder="Enter sentence here...", label="Input Text"), + gr.HighlightedText(label="Output Text"), + examples=examples, + allow_flagging="never") + + if __name__ == "__main__": + try: + demo.launch(debug=True) + except Exception: + demo.launch(share=True, debug=True) + # if you are launching remotely, specify server_name and server_port + # demo.launch(server_name='your server name', server_port='server port in int') + # Read more in the docs: https://gradio.app/docs/ + + +.. parsed-literal:: + + + Thanks for being a Gradio user! If you have questions or feedback, please join our Discord server and chat with us: https://discord.gg/feTf9x3ZSB + Running on local URL: http://127.0.0.1:7860 + + To create a public link, set `share=True` in `launch()`. + + + +.. raw:: html + +
+ + +.. parsed-literal:: + + Keyboard interruption in main thread... closing server. + + +Compare the Original and Quantized Models `⇑ <#top>`__ +############################################################################################################################### + + +Compare the original +`distilbert-base-cased-finetuned-conll03-english `__ +model with quantized and converted to OpenVINO IR format models to see +the difference. + +Compare performance `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +.. code:: ipython3 + + ner_pipeline_original = pipeline("token-classification", model=model, tokenizer=tokenizer) + +.. code:: ipython3 + + import time + import numpy as np + + def calc_perf(ner_pipeline): + inference_times = [] + + for data in calibration_dataset: + text = ' '.join(data['tokens']) + start = time.perf_counter() + ner_pipeline(text) + end = time.perf_counter() + inference_times.append(end - start) + + return np.median(inference_times) + + + print( + f"Median inference time of quantized model: {calc_perf(ner_pipeline_optimized)} " + ) + + print( + f"Median inference time of original model: {calc_perf(ner_pipeline_original)} " + ) + + +.. parsed-literal:: + + Median inference time of quantized model: 0.008888308017048985 + + +Compare size of the models `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +.. code:: ipython3 + + from pathlib import Path + + print(f'Size of original model in Bytes is {Path(original_ner_model_dir, "pytorch_model.bin").stat().st_size}') + print(f'Size of quantized model in Bytes is {Path(quantized_ner_model_dir, "openvino_model.bin").stat().st_size}') diff --git a/docs/notebooks/245-typo-detector-with-output.rst b/docs/notebooks/245-typo-detector-with-output.rst new file mode 100644 index 00000000000..b5e609416e8 --- /dev/null +++ b/docs/notebooks/245-typo-detector-with-output.rst @@ -0,0 +1,593 @@ +Typo Detector with OpenVINO™ +============================ + +Typo detection in AI is a process of identifying and correcting +typographical errors in text data using machine learning algorithms. The +goal of typo detection is to improve the accuracy, readability, and +usability of text by identifying and indicating mistakes made during the +writing process. To detect typos, AI-based typo detectors use various +techniques, such as natural language processing (NLP), machine learning +(ML), and deep learning (DL). + +A typo detector takes a sentence as an input and identify all +typographical errors such as misspellings and homophone errors. + +This tutorial provides how to use the `Typo +Detector `__ +from the `Hugging Face +Transformers `__ library +in the OpenVINO environment to perform the above task. + +The model detects typos in a given text with a high accuracy, +performances of which are listed below, - Precision score of 0.9923 - +Recall score of 0.9859 - f1-score of 0.9891 + +`Source for above +metrics `__ + +These metrics indicate that the model can correctly identify a high +proportion of both correct and incorrect text, minimizing both false +positives and false negatives. + +The model has been pretrained on the +`NeuSpell `__ dataset. + +Imports +~~~~~~~ + +.. code:: ipython3 + + from transformers import AutoConfig, AutoTokenizer, AutoModelForTokenClassification, pipeline + from pathlib import Path + import numpy as np + import torch + import re + from typing import List, Dict + import time + + +.. parsed-literal:: + + 2023-08-16 01:01:23.631663: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-08-16 01:01:23.665285: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. + 2023-08-16 01:01:24.208556: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + + +Methods +~~~~~~~ + +The notebook provides two methods to run the inference of typo detector +with OpenVINO runtime, so that you can experience both calling the API +of Optimum with OpenVINO Runtime included, and loading models in other +frameworks, converting them to OpenVINO IR format, and running inference +with OpenVINO Runtime. + +1. Using the `Hugging Face Optimum `__ library +''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''''' + +The Hugging Face Optimum API is a high-level API that allows us to +convert models from the Hugging Face Transformers library to the +OpenVINO™ IR format. Compiled models in OpenVINO IR format can be loaded +using Optimum. Optimum allows the use of optimization on targeted +hardware. + +2. Converting the model to ONNX and then to OpenVINO IR +''''''''''''''''''''''''''''''''''''''''''''''''''''''' + +First the Pytorch model is converted to the ONNX format and then the +`Model +Optimizer `__ +tool will be used to convert to `OpenVINO IR +format `__. This +method provides much more insight to how to set up a pipeline from model +loading to model converting, compiling and running inference with +OpenVINO, so that you could conveniently use OpenVINO to optimize and +accelerate inference for other deep-learning models. The optimization of +targeted hardware is also used here. + +The following table summarizes the major differences between the two +methods + ++-----------------------------------+----------------------------------+ +| Method 1 | Method 2 | ++===================================+==================================+ +| Load models from Optimum, an | Load model from transformers | +| extension of transformers | | ++-----------------------------------+----------------------------------+ +| Load the model in OpenVINO IR | Convert to ONNX and then to | +| format on the fly | OpenVINO IR | ++-----------------------------------+----------------------------------+ +| Load the compiled model by | Compile the OpenVINO IR and run | +| default | inference with OpenVINO Runtime | ++-----------------------------------+----------------------------------+ +| Pipeline is created to run | Manually run inference. | +| inference with OpenVINO Runtime | | ++-----------------------------------+----------------------------------+ + +Select inference device +~~~~~~~~~~~~~~~~~~~~~~~ + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + from openvino.runtime import Core + + core = Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +1. Hugging Face Optimum Intel library +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +For this method, we need to install the +``Hugging Face Optimum Intel library`` accelerated by OpenVINO +integration. + +Optimum Intel can be used to load optimized models from the `Hugging +Face Hub `__ and +create pipelines to run an inference with OpenVINO Runtime using Hugging +Face APIs. The Optimum Inference models are API compatible with Hugging +Face Transformers models. This means we need just replace +``AutoModelForXxx`` class with the corresponding ``OVModelForXxx`` +class. + +.. code:: ipython3 + + !pip install -q "diffusers>=0.17.1" "openvino-dev>=2023.0.0" "nncf>=2.5.0" "gradio" "onnx>=1.11.0" "onnxruntime>=1.14.0" "optimum-intel>=1.9.1" "transformers>=4.31.0" + + +.. parsed-literal:: + + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + + +Import required model class + +.. code:: ipython3 + + from optimum.intel.openvino import OVModelForTokenClassification + + +.. parsed-literal:: + + INFO:nncf:NNCF initialized successfully. Supported frameworks detected: torch, tensorflow, onnx, openvino + + +.. parsed-literal:: + + No CUDA runtime is found, using CUDA_HOME='/usr/local/cuda' + + +Load the model +'''''''''''''' + +From the ``OVModelForTokenCLassification`` class we will import the +relevant pre-trained model. To load a Transformers model and convert it +to the OpenVINO format on-the-fly, we set ``export=True`` when loading +your model. + +.. code:: ipython3 + + # The pretrained model we are using + model_id = "m3hrdadfi/typo-detector-distilbert-en" + + model_dir = Path("optimum_model") + + # Save the model to the path if not existing + if model_dir.exists(): + model = OVModelForTokenClassification.from_pretrained(model_dir, device=device.value) + else: + model = OVModelForTokenClassification.from_pretrained(model_id, export=True, device=device.value) + model.save_pretrained(model_dir) + + +.. code:: + + Framework not specified. Using pt to export to ONNX. + Using framework PyTorch: 1.13.1+cpu + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/nncf/torch/dynamic_graph/wrappers.py:74: TracerWarning: torch.tensor results are registered as constants in the trace. You can safely ignore this warning if you use this function to create tensors out of constant variables that would be the same every time you call this function. In any other case, this might cause the trace to be incorrect. + op1 = operator(*args, **kwargs) + Compiling the model... + Set CACHE_DIR to /tmp/tmpmevydbbe/model_cache + + +Load the tokenizer +'''''''''''''''''' + +Text Preprocessing cleans the text-based input data so it can be fed +into the model. Tokenization splits paragraphs and sentences into +smaller units that can be more easily assigned meaning. It involves +cleaning the data and assigning tokens or IDs to the words, so they are +represented in a vector space where similar words have similar vectors. +This helps the model understand the context of a sentence. We’re making +use of an +`AutoTokenizer `__ +from Hugging Face, which is essentially a pretrained tokenizer. + +.. code:: ipython3 + + tokenizer = AutoTokenizer.from_pretrained(model_id) + +Then we use the inference pipeline for ``token-classification`` task. +You can find more information about usage Hugging Face inference +pipelines in this +`tutorial `__ + +.. code:: ipython3 + + nlp = pipeline('token-classification', model=model, tokenizer=tokenizer, aggregation_strategy="average") + +Function to find typos in a sentence and write them to the terminal + +.. code:: ipython3 + + def show_typos(sentence: str): + """ + Detect typos from the given sentence. + Writes both the original input and typo-tagged version to the terminal. + + Arguments: + sentence -- Sentence to be evaluated (string) + """ + + typos = [sentence[r["start"]: r["end"]] for r in nlp(sentence)] + + detected = sentence + for typo in typos: + detected = detected.replace(typo, f'{typo}') + + print("[Input]: ", sentence) + print("[Detected]: ", detected) + print("-" * 130) + +Let’s run a demo using the Hugging Face Optimum API. + +.. code:: ipython3 + + sentences = [ + "He had also stgruggled with addiction during his time in Congress .", + "The review thoroughla assessed all aspects of JLENS SuR and CPG esign maturit and confidence .", + "Letterma also apologized two his staff for the satyation .", + "Vincent Jay had earlier won France 's first gold in gthe 10km biathlon sprint .", + "It is left to the directors to figure out hpw to bring the stry across to tye audience .", + "I wnet to the park yestreday to play foorball with my fiends, but it statred to rain very hevaily and we had to stop.", + "My faorite restuarant servs the best spahgetti in the town, but they are always so buzy that you have to make a resrvation in advnace.", + "I was goig to watch a mvoie on Netflx last night, but the straming was so slow that I decided to cancled my subscrpition.", + "My freind and I went campign in the forest last weekend and saw a beutiful sunst that was so amzing it took our breth away.", + "I have been stuying for my math exam all week, but I'm stil not very confidet that I will pass it, because there are so many formuals to remeber." + ] + + start = time.time() + + for sentence in sentences: + show_typos(sentence) + + print(f"Time elapsed: {time.time() - start}") + + +.. parsed-literal:: + + [Input]: He had also stgruggled with addiction during his time in Congress . + [Detected]: He had also stgruggled with addiction during his time in Congress . + ---------------------------------------------------------------------------------------------------------------------------------- + [Input]: The review thoroughla assessed all aspects of JLENS SuR and CPG esign maturit and confidence . + [Detected]: The review thoroughla assessed all aspects of JLENS SuR and CPG esign maturit and confidence . + ---------------------------------------------------------------------------------------------------------------------------------- + [Input]: Letterma also apologized two his staff for the satyation . + [Detected]: Letterma also apologized two his staff for the satyation . + ---------------------------------------------------------------------------------------------------------------------------------- + [Input]: Vincent Jay had earlier won France 's first gold in gthe 10km biathlon sprint . + [Detected]: Vincent Jay had earlier won France 's first gold in gthe 10km biathlon sprint . + ---------------------------------------------------------------------------------------------------------------------------------- + [Input]: It is left to the directors to figure out hpw to bring the stry across to tye audience . + [Detected]: It is left to the directors to figure out hpw to bring the stry across to tye audience . + ---------------------------------------------------------------------------------------------------------------------------------- + [Input]: I wnet to the park yestreday to play foorball with my fiends, but it statred to rain very hevaily and we had to stop. + [Detected]: I wnet to the park yestreday to play foorball with my fiends, but it statred to rain very hevaily and we had to stop. + ---------------------------------------------------------------------------------------------------------------------------------- + [Input]: My faorite restuarant servs the best spahgetti in the town, but they are always so buzy that you have to make a resrvation in advnace. + [Detected]: My faorite restuarant servs the best spahgetti in the town, but they are always so buzy that you have to make a resrvation in advnace. + ---------------------------------------------------------------------------------------------------------------------------------- + [Input]: I was goig to watch a mvoie on Netflx last night, but the straming was so slow that I decided to cancled my subscrpition. + [Detected]: I was goig to watch a mvoie on Netflx last night, but the straming was so slow that I decided to cancled my subscrpition. + ---------------------------------------------------------------------------------------------------------------------------------- + [Input]: My freind and I went campign in the forest last weekend and saw a beutiful sunst that was so amzing it took our breth away. + [Detected]: My freind and I went campign in the forest last weekend and saw a beutiful sunst that was so amzing it took our breth away. + ---------------------------------------------------------------------------------------------------------------------------------- + [Input]: I have been stuying for my math exam all week, but I'm stil not very confidet that I will pass it, because there are so many formuals to remeber. + [Detected]: I have been stuying for my math exam all week, but I'm stil not very confidet that I will pass it, because there are so many formuals to remeber. + ---------------------------------------------------------------------------------------------------------------------------------- + Time elapsed: 0.20883584022521973 + + +2. Converting the model to ONNX and then to OpenVINO IR +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +Load the Pytorch model +'''''''''''''''''''''' + +Use the ``AutoModelForTokenClassification`` class to load the pretrained +pytorch model. + +.. code:: ipython3 + + model_id = "m3hrdadfi/typo-detector-distilbert-en" + model_dir = Path("pytorch_model") + + tokenizer = AutoTokenizer.from_pretrained(model_id) + config = AutoConfig.from_pretrained(model_id) + + # Save the model to the path if not existing + if model_dir.exists(): + model = AutoModelForTokenClassification.from_pretrained(model_dir) + else: + model = AutoModelForTokenClassification.from_pretrained(model_id, config=config) + model.save_pretrained(model_dir) + +Converting to `ONNX `__ +''''''''''''''''''''''''''''''''''''''''' + +``ONNX`` is an open format built to represent machine learning models. +ONNX defines a common set of operators - the building blocks of machine +learning and deep learning models - and a common file format to enable +AI developers to use models with a variety of frameworks, tools, +runtimes, and compilers. We need to convert our model from PyTorch to +ONNX. In order to perform the operation, we use the torch.onnx.export +function to `convert a Hugging Face +model `__ +to its respective ONNX format. + +.. code:: ipython3 + + onnx_model = "typo_detect.onnx" + + onnx_model_path = Path(model_dir) / onnx_model + + dummy_model_input = tokenizer("This is a sample", return_tensors="pt") + + torch.onnx.export( + model, + tuple(dummy_model_input.values()), + f=onnx_model_path, + input_names=['input_ids', 'attention_mask'], + output_names=['logits'], + dynamic_axes={'input_ids': {0: 'batch_size', 1: 'sequence'}, + 'attention_mask': {0: 'batch_size', 1: 'sequence'}, + 'logits': {0: 'batch_size', 1: 'sequence'}}, + ) + +Model Optimizer +''''''''''''''' + +`Model +Optimizer `__ +is a cross-platform command-line tool that facilitates the transition +between training and deployment environments, performs static model +analysis, and adjusts deep learning models for optimal execution on +end-point target devices. Model Optimizer converts the model to the +OpenVINO Intermediate Representation format (IR), which you can infer +later with `OpenVINO +runtime `__. + +.. code:: ipython3 + + from openvino.tools.mo import convert_model + + ov_model = convert_model(onnx_model_path) + +Inference +''''''''' + +OpenVINO™ Runtime Python API is used to compile the model in OpenVINO IR +format. The +`Core `__ +class from the ``openvino.runtime`` module is imported first. This class +provides access to the OpenVINO Runtime API. The ``core`` object, which +is an instance of the ``Core`` class, represents the API and it is used +to compile the model. The output layer is extracted from the compiled +model as it is needed for inference. + +.. code:: ipython3 + + from openvino.runtime import Core + + core = Core() + compiled_model = core.compile_model(ov_model, device.value) + output_layer = compiled_model.output(0) + +Helper Functions +~~~~~~~~~~~~~~~~ + +.. code:: ipython3 + + def token_to_words(tokens: List[str]) -> Dict[str, int]: + """ + Maps the list of tokens to words in the original text. + Built on the feature that tokens starting with '##' is attached to the previous token as tokens derived from the same word. + + Arguments: + tokens -- List of tokens + + Returns: + map_to_words -- Dictionary mapping tokens to words in original text + """ + + word_count = -1 + map_to_words = {} + for token in tokens: + if token.startswith('##'): + map_to_words[token] = word_count + continue + word_count += 1 + map_to_words[token] = word_count + return map_to_words + +.. code:: ipython3 + + def infer(input_text: str) -> Dict[np.ndarray, np.ndarray]: + """ + Creating a generic inference function to read the input and infer the result + + Arguments: + input_text -- The text to be infered (String) + + Returns: + result -- Resulting list from inference + """ + + tokens = tokenizer( + input_text, + return_tensors="np", + ) + inputs = dict(tokens) + result = compiled_model(inputs)[output_layer] + return result + +.. code:: ipython3 + + def get_typo_indexes(result: Dict[np.ndarray, np.ndarray], map_to_words: Dict[str, int], tokens: List[str]) -> List[int]: + """ + Given results from the inference and tokens-map-to-words, identifies the indexes of the words with typos. + + Arguments: + result -- Result from inference (tensor) + map_to_words -- Dictionary mapping tokens to words (Dictionary) + + Results: + wrong_words -- List of indexes of words with typos + """ + + wrong_words = [] + c = 0 + result_list = result[0][1:-1] + for i in result_list: + prob = np.argmax(i) + if prob == 1: + if map_to_words[tokens[c]] not in wrong_words: + wrong_words.append(map_to_words[tokens[c]]) + c += 1 + return wrong_words + +.. code:: ipython3 + + def sentence_split(sentence: str) -> List[str]: + """ + Split the sentence into words and characters + + Arguments: + sentence - Sentence to be split (string) + + Returns: + splitted -- List of words and characters + """ + + splitted = re.split("([',. ])",sentence) + splitted = [x for x in splitted if x != " " and x != ""] + return splitted + +.. code:: ipython3 + + def show_typos(sentence: str): + """ + Detect typos from the given sentence. + Writes both the original input and typo-tagged version to the terminal. + + Arguments: + sentence -- Sentence to be evaluated (string) + """ + + tokens = tokenizer.tokenize(sentence) + map_to_words = token_to_words(tokens) + result = infer(sentence) + typo_indexes = get_typo_indexes(result,map_to_words, tokens) + + sentence_words = sentence_split(sentence) + + typos = [sentence_words[i] for i in typo_indexes] + + detected = sentence + for typo in typos: + detected = detected.replace(typo, f'{typo}') + + print(" [Input]: ", sentence) + print("[Detected]: ", detected) + print("-" * 130) + +Let’s run a demo using the converted OpenVINO IR model. + +.. code:: ipython3 + + sentences = [ + "He had also stgruggled with addiction during his time in Congress .", + "The review thoroughla assessed all aspects of JLENS SuR and CPG esign maturit and confidence .", + "Letterma also apologized two his staff for the satyation .", + "Vincent Jay had earlier won France 's first gold in gthe 10km biathlon sprint .", + "It is left to the directors to figure out hpw to bring the stry across to tye audience .", + "I wnet to the park yestreday to play foorball with my fiends, but it statred to rain very hevaily and we had to stop.", + "My faorite restuarant servs the best spahgetti in the town, but they are always so buzy that you have to make a resrvation in advnace.", + "I was goig to watch a mvoie on Netflx last night, but the straming was so slow that I decided to cancled my subscrpition.", + "My freind and I went campign in the forest last weekend and saw a beutiful sunst that was so amzing it took our breth away.", + "I have been stuying for my math exam all week, but I'm stil not very confidet that I will pass it, because there are so many formuals to remeber." + ] + + start = time.time() + + for sentence in sentences: + show_typos(sentence) + + print(f"Time elapsed: {time.time() - start}") + + +.. parsed-literal:: + + [Input]: He had also stgruggled with addiction during his time in Congress . + [Detected]: He had also stgruggled with addiction during his time in Congress . + ---------------------------------------------------------------------------------------------------------------------------------- + [Input]: The review thoroughla assessed all aspects of JLENS SuR and CPG esign maturit and confidence . + [Detected]: The review thoroughla assessed all aspects of JLENS SuR and CPG esign maturit and confidence . + ---------------------------------------------------------------------------------------------------------------------------------- + [Input]: Letterma also apologized two his staff for the satyation . + [Detected]: Letterma also apologized two his staff for the satyation . + ---------------------------------------------------------------------------------------------------------------------------------- + [Input]: Vincent Jay had earlier won France 's first gold in gthe 10km biathlon sprint . + [Detected]: Vincent Jay had earlier won France 's first gold in gthe 10km biathlon sprint . + ---------------------------------------------------------------------------------------------------------------------------------- + [Input]: It is left to the directors to figure out hpw to bring the stry across to tye audience . + [Detected]: It is left to the directors to figure out hpw to bring the stry across to tye audience . + ---------------------------------------------------------------------------------------------------------------------------------- + [Input]: I wnet to the park yestreday to play foorball with my fiends, but it statred to rain very hevaily and we had to stop. + [Detected]: I wnet to the park yestreday to play foorball with my fiends, but it statred to rain very hevaily and we had to stop. + ---------------------------------------------------------------------------------------------------------------------------------- + [Input]: My faorite restuarant servs the best spahgetti in the town, but they are always so buzy that you have to make a resrvation in advnace. + [Detected]: My faorite restuarant servs the best spahgetti in the town, but they are always so buzy that you have to make a resrvation in advnace. + ---------------------------------------------------------------------------------------------------------------------------------- + [Input]: I was goig to watch a mvoie on Netflx last night, but the straming was so slow that I decided to cancled my subscrpition. + [Detected]: I was goig to watch a mvoie on Netflx last night, but the straming was so slow that I decided to cancled my subscrpition. + ---------------------------------------------------------------------------------------------------------------------------------- + [Input]: My freind and I went campign in the forest last weekend and saw a beutiful sunst that was so amzing it took our breth away. + [Detected]: My freind and I went campign in the forest last weekend and saw a beutiful sunst that was so amzing it took our breth away. + ---------------------------------------------------------------------------------------------------------------------------------- + [Input]: I have been stuying for my math exam all week, but I'm stil not very confidet that I will pass it, because there are so many formuals to remeber. + [Detected]: I have been stuying for my math exam all week, but I'm stil not very confidet that I will pass it, because there are so many formuals to remeber. + ---------------------------------------------------------------------------------------------------------------------------------- + Time elapsed: 0.1267991065979004 + diff --git a/docs/notebooks/246-depth-estimation-videpth-with-output.rst b/docs/notebooks/246-depth-estimation-videpth-with-output.rst new file mode 100644 index 00000000000..b515d590807 --- /dev/null +++ b/docs/notebooks/246-depth-estimation-videpth-with-output.rst @@ -0,0 +1,1027 @@ +Monocular Visual-Inertial Depth Estimation using OpenVINO™ +========================================================== + +.. raw:: html + +

+ +.. raw:: html + +

+ +The overall methodology. Diagram taken from the VI-Depth repository. + +.. raw:: html + +
+ +.. raw:: html + +

+ +A visual-inertial depth estimation pipeline that integrates monocular +depth estimation and visual-inertial odometry to produce dense depth +estimates with metric scale has been presented by the authors. The +approach consists of three stages: + +1. input processing, where RGB and inertial measurement unit (IMU) data + feed into monocular depth estimation alongside visual-inertial + odometry, +2. global scale and shift alignment, where monocular depth estimates are + fitted to sparse depth from visual inertial odometry (VIO) in a + least-squares manner and +3. learning-based dense scale alignment, where globally-aligned depth is + locally realigned using a dense scale map regressed by the + ScaleMapLearner (SML). + +The images at the bottom in the diagram above illustrate a Visual +Odometry with Inertial and Depth (VOID) sample being processed through +our pipeline; from left to right: the input RGB, ground truth depth, +sparse depth from VIO, globally-aligned depth, scale map scaffolding, +dense scale map regressed by SML, final depth output. + +.. raw:: html + +

+ +.. raw:: html + +

+ +An illustration of VOID samples being processed by the image pipeline. +Image taken from the VI-Depth repository. + +.. raw:: html + +
+ +.. raw:: html + +

+ +We will be consulting the `VI-Depth +repository `__ for the +pre-processing, model transformations and basic utility code. A part of +it has already been kept as it is in the `utils `__ directory. At +the same time we will learn how to perform `model +conversion `__ +for converting a model in a different format to the standard OpenVINO™ +IR model representation *via* another format. + +Imports +~~~~~~~ + +.. code:: ipython3 + + # Import sys beforehand to inform of Python version <= 3.7 not being supported + import sys + + if sys.version_info.minor < 8: + print('Python3.7 is not supported. Some features might not work as expected') + + # Download the correct version of the PyTorch deep learning library associated with image models + # alongside the lightning module + !pip install -q lightning timm==0.6.12 + + +.. parsed-literal:: + + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + ERROR: pip's dependency resolver does not currently take into account all the packages that are installed. This behaviour is the source of the following dependency conflicts. + onnx 1.14.0 requires protobuf>=3.20.2, but you have protobuf 3.20.1 which is incompatible. + paddlepaddle 2.5.0rc0 requires protobuf>=3.20.2; platform_system != "Windows", but you have protobuf 3.20.1 which is incompatible. + tensorflow 2.12.0 requires protobuf!=4.21.0,!=4.21.1,!=4.21.2,!=4.21.3,!=4.21.4,!=4.21.5,<5.0.0dev,>=3.20.3, but you have protobuf 3.20.1 which is incompatible. + + +.. code:: ipython3 + + import matplotlib.pyplot as plt + import matplotlib.image as mpimg + import numpy as np + import openvino + import torch + import torchvision + from openvino.runtime import Core + from pathlib import Path + from shutil import rmtree + from typing import Optional, Tuple + + sys.path.append('../utils') + from notebook_utils import download_file + + sys.path.append('vi_depth_utils') + import data_loader + import modules.midas.transforms as transforms + import modules.midas.utils as utils + from modules.estimator import LeastSquaresEstimator + from modules.interpolator import Interpolator2D + from modules.midas.midas_net_custom import MidasNet_small_videpth + +.. code:: ipython3 + + # Ability to display images inline + %matplotlib inline + +Loading models and checkpoints +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +The complete pipeline here requires only two models: one for depth +estimation and a ScaleMapLearner model which is responsible for +regressing a dense scale map. The table of models which has been given +in the original `VI-Depth repo `__ +has been presented as it is for the users to download from. +`VOID `__ is the name of the +original dataset from on which these models have been trained. The +numbers after the word **VOID** represent the checkpoint in the model +obtained after training samples for sparse dense maps corresponding to +:math:`150`, :math:`500` and :math:`1500` levels in the density map. +Just *right-click* on any of the highlighted links and click on “Copy +link address”. We shall use this link in the next cell to download the +ScaleMapLearner model. *Interestingly*, the ScaleMapLearner decides the +depth prediction model as you will see. + + ++------------------+---------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------+ +| Depth Predictor | SML on VOID 150 | SML on VOID 500 | SML on VOID 1500 | ++==================+=================================================================================================================================+==================================================================================================================================+===================================================================================================================================+ +| DPT-BEiT-Large | `model `__ | `model `__ | `model `__ | ++------------------+---------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------+ +| DPT-SwinV2-Large | `model `__ | `model `__ | `model `__ | ++------------------+---------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------+ +| DPT-Large | `model `__ | `model `__ | `model `__ | ++------------------+---------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------+ +| DPT-Hybrid | `model `__ \* | `model `__ | `model `__ | ++------------------+---------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------+ +| DPT-SwinV2-Tiny | `model `__ | `model `__ | `model `__ | ++------------------+---------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------+ +| DPT-LeViT | `model `__ | `model `__ | `model `__ | ++------------------+---------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------+ +| MiDaS-small | `model `__ | `model `__ | `model `__ | ++------------------+---------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------+ + + +\* Also available with pre-training on TartanAir: +`model `__ + +.. code:: ipython3 + + # Base directory in which models would be stored as a pathlib.Path variable + MODEL_DIR = Path('model') + + # Mapping between depth predictors and the corresponding scale map learners + PREDICTOR_MODEL_MAP = {'dpt_beit_large_512': 'DPT_BEiT_L_512', + 'dpt_swin2_large_384': 'DPT_SwinV2_L_384', + 'dpt_large': 'DPT_Large', + 'dpt_hybrid': 'DPT_Hybrid', + 'dpt_swin2_tiny_256': 'DPT_SwinV2_T_256', + 'dpt_levit_224': 'DPT_LeViT_224', + 'midas_small': 'MiDaS_small'} + +.. code:: ipython3 + + # Create the model directory adjacent to the notebook and suppress errors if the directory already exists + MODEL_DIR.mkdir(exist_ok=True) + + # Here we will be downloading the SML model corresponding to the MiDaS-small depth predictor for + # the checkpoint captured after training on 1500 points of the density level. Suppress errors if the file already exists + download_file('https://github.com/isl-org/VI-Depth/releases/download/v1/sml_model.dpredictor.midas_small.nsamples.1500.ckpt', directory=MODEL_DIR, silent=True) + + # Take a note of the samples. It would be of major use later on + NSAMPLES = 1500 + + + +.. parsed-literal:: + + model/sml_model.dpredictor.midas_small.nsamples.1500.ckpt: 0%| | 0.00/208M [00:00 str: + """ + Download a model from the pre-validated 'isl-org/MiDaS:2.1' set of releases on the GitHub repo + while simultaneously trusting the repo permanently + + :param: depth_predictor: Any depth estimation model amongst the ones given at https://github.com/isl-org/VI-Depth#setup + :param: remote_repo: The remote GitHub repo from where the models will be downloaded + :returns: A PyTorch model callable + """ + + # Workaround for avoiding rate limit errors + torch.hub._validate_not_a_forked_repo = lambda a, b, c: True + + return torch.hub.load(remote_repo, PREDICTOR_MODEL_MAP[depth_predictor], skip_validation=True, trust_repo=True) + +.. code:: ipython3 + + # Execute the above function so as to download the MiDaS-small model + # and get the output of the model callable in return + depth_model = get_model_for_predictor('midas_small') + + +.. parsed-literal:: + + Downloading: "https://github.com/intel-isl/MiDaS/zipball/master" to model/master.zip + + +.. parsed-literal:: + + Loading weights: None + + +.. parsed-literal:: + + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/hub.py:267: UserWarning: You are about to download and run code from an untrusted repository. In a future release, this won't be allowed. To add the repository to your trusted list, change the command to {calling_fn}(..., trust_repo=False) and a command prompt will appear asking for an explicit confirmation of trust, or load(..., trust_repo=True), which will assume that the prompt is to be answered with 'yes'. You can also use load(..., trust_repo='check') which will only prompt for confirmation if the repo is not already trusted. This will eventually be the default behaviour + warnings.warn( + Downloading: "https://github.com/rwightman/gen-efficientnet-pytorch/zipball/master" to model/master.zip + Downloading: "https://github.com/rwightman/pytorch-image-models/releases/download/v0.1-weights/tf_efficientnet_lite3-b733e338.pth" to model/checkpoints/tf_efficientnet_lite3-b733e338.pth + Downloading: "https://github.com/isl-org/MiDaS/releases/download/v2_1/midas_v21_small_256.pt" to model/checkpoints/midas_v21_small_256.pt + + + +.. parsed-literal:: + + 0%| | 0.00/81.8M [00:00`__ +downloads a lot of unnecessary files. We shall move remove the +unnecessary directories and files which were created during the download +process. + +.. code:: ipython3 + + # Remove unnecessary directories and files and suppress errors(if any) + rmtree(path=str(MODEL_DIR / 'intel-isl_MiDaS_master'), ignore_errors=True) + rmtree(path=str(MODEL_DIR / 'rwightman_gen-efficientnet-pytorch_master'), ignore_errors=True) + rmtree(path=str(MODEL_DIR / 'checkpoints'), ignore_errors=True) + + # Check for the existence of the trusted list file and then remove + list_file = Path(MODEL_DIR / 'trusted_list') + if list_file.is_file(): + list_file.unlink() + +Transformation of models +~~~~~~~~~~~~~~~~~~~~~~~~ + +Each of the models need an appropriate transformation which can be +invoked by the ``get_model_transforms`` function. It needs only the +``depth_predictor`` parameter and ``NSAMPLES`` defined above to work. +The reason being that the ``ScaleMapLearner`` and the depth estimation +model are always in direct correspondence with each other. + +.. code:: ipython3 + + # Define important custom types + type_transform_compose = torchvision.transforms.transforms.Compose + type_compiled_model = openvino.runtime.ie_api.CompiledModel + +.. code:: ipython3 + + def get_model_transforms(depth_predictor: str, nsamples: int) -> Tuple[type_transform_compose, type_transform_compose]: + """ + Construct the transformation of the depth prediction model and the + associated ScaleMapLearner model + + :param: depth_predictor: Any depth estimation model amongst the ones given at https://github.com/isl-org/VI-Depth#setup + :param: nsamples: The no. of density levels for the depth map + :returns: The transformed models as the resut of torchvision.transforms.Compose operations + """ + model_transforms = transforms.get_transforms(depth_predictor, "void", str(nsamples)) + return model_transforms['depth_model'], model_transforms['sml_model'] + +.. code:: ipython3 + + # Obtain transforms of both the models here + depth_model_transform, scale_map_learner_transform = get_model_transforms(depth_predictor='midas_small', + nsamples=NSAMPLES) + +Dummy input creation +^^^^^^^^^^^^^^^^^^^^ + +Dummy inputs are necessary for `PyTorch to +ONNX `__ +conversion. Although +`torch.onnx.export `__ +accepts any dummy input for a single pass through the model and thereby +enabling model conversion, the pre-processing required for the actual +inputs later at inference using compiled models would be substantial. So +we have decided that even dummy inputs should go through the proper +transformation process so that the reader gets the idea of a +*transformed image* being compiled by a *transformed model*. + +Also note down the width and height of the image which would be used +multiple times later. Do note that this is constant throughout the +dataset + +.. code:: ipython3 + + IMAGE_H, IMAGE_W = 480, 640 + + # Although you can always verify the same by uncommenting and running + # the following lines + # img = cv2.imread('data/image/dummy_img.png') + # print(img.shape) + +.. code:: ipython3 + + # Base directory in which data would be stored as a pathlib.Path variable + DATA_DIR = Path('data') + + # Create the data directory tree adjacent to the notebook and suppress errors if the directory already exists + # Create a directory each for the images and their corresponding depth maps + DATA_DIR.mkdir(exist_ok=True) + Path(DATA_DIR / 'image').mkdir(exist_ok=True) + Path(DATA_DIR / 'sparse_depth').mkdir(exist_ok=True) + + # Download the dummy image and its depth scale (take a note of the image hashes for possible later use) + # On the fly download is being done to avoid unnecessary memory/data load during testing and + # creation of PRs + download_file('https://user-images.githubusercontent.com/22426058/254174385-161b9f0e-5991-4308-ba89-d81bc02bcb7c.png', filename='dummy_img.png', directory=Path(DATA_DIR / 'image'), silent=True) + download_file('https://user-images.githubusercontent.com/22426058/254174398-8c71c59f-0adf-43c6-ad13-c04431e02349.png', filename='dummy_depth.png', directory=Path(DATA_DIR / 'sparse_depth'), silent=True) + + # Load the dummy image and its depth scale + dummy_input = data_loader.load_input_image('data/image/dummy_img.png') + dummy_depth = data_loader.load_sparse_depth('data/sparse_depth/dummy_depth.png') + + + +.. parsed-literal:: + + data/image/dummy_img.png: 0%| | 0.00/328k [00:00 torch.Tensor: + """ + Transform the input_image for processing by a PyTorch depth estimation model + + :param: input_image: The input image obtained as a result of data_loader.load_input_image + :param: depth_model_transform: The transformed depth model + :param: device: The device on which the image would be allocated + :returns: The transformed image suitable to be used as an input to the depth estimation model + """ + input_height, input_width = np.shape(input_image)[:2] + + sample = {'image' : input_image} + sample = depth_model_transform(sample) + im = sample['image'].to(device) + return im.unsqueeze(0) + +.. code:: ipython3 + + # Transform the dummy input image for the depth model + transformed_dummy_image = transform_image_for_depth(input_image=dummy_input, depth_model_transform=depth_model_transform) + +Conversion of depth model to OpenVINO™ IR format +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +The OpenVINO™ toolkit doesn’t provide any direct method of converting +PyTorch models to the intermediate representation format. To have a +depth estimation model in the OpenVINO™ IR format and then compile it, +we shall follow the following steps: + +1. Use the ``depth_model`` callable to our advantage from the *Loading + models and checkpoints* stage. +2. Export the model to ``.onnx`` format using the transformed dummy + input created earlier. +3. Use the serialize function from OpenVINO to create equivalent + ``.xml`` and ``.bin`` files and obtain compiled models in the same + step. Alternatively serialization procedure may be avoided and + compiled model may be obtained by directly using OpenVINO’s + ``compile`` function. + +.. code:: ipython3 + + # Evaluate the model to switch some operations from training mode to inference. + depth_model.eval() + + # Call the export function via the transformed dummy image obtained from the + # previous step. It is absolutely not a case of worry if multiple warnings pop up + # in this step. They can be safely ignored. + torch.onnx.export(model=depth_model, args=(transformed_dummy_image, ), f=str(MODEL_DIR / 'depth_model.onnx')) + + +.. parsed-literal:: + + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/246-depth-estimation-videpth/model/rwightman_gen-efficientnet-pytorch_master/geffnet/conv2d_layers.py:47: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/_internal/jit_utils.py:258: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + _C._jit_pass_onnx_node_shape_type_inference(node, params_dict, opset_version) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) + _C._jit_pass_onnx_graph_shape_type_inference( + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + _C._jit_pass_onnx_graph_shape_type_inference( + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) + _C._jit_pass_onnx_graph_shape_type_inference( + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + _C._jit_pass_onnx_graph_shape_type_inference( + + +Select inference device +''''''''''''''''''''''' + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + core = Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +Compilation of depth model +'''''''''''''''''''''''''' + +Now we can go ahead and compile our depth models from the ``.onnx`` file +path. We will not perform serialization because we don’t plan to re-read +the file again within this tutorial. Therefore we will use the compiled +depth estimation model as it is. + +.. code:: ipython3 + + # Initialize OpenVINO Runtime. + core = Core() + depth_model = core.read_model(MODEL_DIR / 'depth_model.onnx') + compiled_depth_model = core.compile_model(model=depth_model, device_name=device.value) + +.. code:: ipython3 + + def run_depth_model(input_image_h: int, input_image_w: int, + transformed_image: torch.Tensor, compiled_depth_model: type_compiled_model) -> np.ndarray: + + """ + Run the compiled_depth_model on the transformed_image of dimensions + input_image_w x input_image_h + + :param: input_image_h: The height of the input image + :param: input_image_w: The width of the input image + :param: transformed_image: The transformed image suitable to be used as an input to the depth estimation model + :returns: + depth_pred: The depth prediction on the image as an np.ndarray type + + """ + + # Obtain the last output layer separately + output_layer_depth_model = compiled_depth_model.output(0) + + with torch.no_grad(): + # Perform computation like a standard OpenVINO compiled model + depth_pred = torch.from_numpy(compiled_depth_model([transformed_image])[output_layer_depth_model]) + depth_pred = ( + torch.nn.functional.interpolate( + depth_pred.unsqueeze(1), + size=(input_image_h, input_image_w), + mode='bicubic', + align_corners=False, + ) + .squeeze() + .cpu() + .numpy() + ) + + return depth_pred + +.. code:: ipython3 + + # Run the compiled depth model using the dummy input + # It will be used to compute the metrics associated with the ScaleMapLearner model + # and hence obtain a compiled version of the same later + depth_pred_dummy = run_depth_model(input_image_h=IMAGE_H, input_image_w=IMAGE_W, + transformed_image=transformed_dummy_image, compiled_depth_model=compiled_depth_model) + +Computation of scale and shift parameters +''''''''''''''''''''''''''''''''''''''''' + +Computation of these parameters required the depth estimation model +output from the previous step. These are the regression based parameters +the ScaleMapLearner model deals with. An utility function for the +purpose has already been created. + +.. code:: ipython3 + + def compute_global_scale_and_shift(input_sparse_depth: np.ndarray, validity_map: Optional[np.ndarray], + depth_pred: np.ndarray, + min_pred: float = 0.1, max_pred: float = 8.0, + min_depth: float = 0.2, max_depth: float = 5.0) -> Tuple[np.ndarray, np.ndarray]: + + """ + Compute the global scale and shift alignment required for SML model to work on + with the input_sparse_depth map being provided and the depth estimation output depth_pred + being provided with an optional validity_map + + :param: input_sparse_depth: The depth map of the input image + :param: validity_map: An optional depth map associated with the original input image + :param: depth_pred: The depth estimate obtained after running the depth model on the input image + :param: min_pred: Lower bound for predicted depth values + :param: max_pred: Upper bound for predicted depth values + :param: min_depth: Min valid depth when evaluating + :param: max_depth: Max valid depth when evaluating + :returns: + int_depth: The depth estimate for the SML regression model + int_scales: The scale to be used for the SML regression model + + """ + + input_sparse_depth_valid = (input_sparse_depth < max_depth) * (input_sparse_depth > min_depth) + if validity_map is not None: + input_sparse_depth_valid *= validity_map.astype(np.bool) + + input_sparse_depth_valid = input_sparse_depth_valid.astype(bool) + input_sparse_depth[~input_sparse_depth_valid] = np.inf # set invalid depth + input_sparse_depth = 1.0 / input_sparse_depth + + # global scale and shift alignment + GlobalAlignment = LeastSquaresEstimator( + estimate=depth_pred, + target=input_sparse_depth, + valid=input_sparse_depth_valid + ) + GlobalAlignment.compute_scale_and_shift() + GlobalAlignment.apply_scale_and_shift() + GlobalAlignment.clamp_min_max(clamp_min=min_pred, clamp_max=max_pred) + int_depth = GlobalAlignment.output.astype(np.float32) + + # interpolation of scale map + assert (np.sum(input_sparse_depth_valid) >= 3), 'not enough valid sparse points' + ScaleMapInterpolator = Interpolator2D( + pred_inv=int_depth, + sparse_depth_inv=input_sparse_depth, + valid=input_sparse_depth_valid, + ) + ScaleMapInterpolator.generate_interpolated_scale_map( + interpolate_method='linear', + fill_corners=False + ) + + int_scales = ScaleMapInterpolator.interpolated_scale_map.astype(np.float32) + int_scales = utils.normalize_unit_range(int_scales) + + return int_depth, int_scales + +.. code:: ipython3 + + # Call the function on the dummy depth map we loaded in the dummy_depth variable + # with all default settings and store in appropriate variables + d_depth, d_scales = compute_global_scale_and_shift(input_sparse_depth=dummy_depth, validity_map=None, depth_pred=depth_pred_dummy) + +.. code:: ipython3 + + def transform_image_for_depth_scale(input_image: np.ndarray, scale_map_learner_transform: type_transform_compose, + int_depth: np.ndarray, int_scales: np.ndarray, + device: torch.device = 'cpu') -> Tuple[torch.Tensor, torch.Tensor]: + """ + Transform the input_image for processing by a PyTorch SML model + + :param: input_image: The input image obtained as a result of data_loader.load_input_image + :param: scale_map_learner_transform: The transformed SML model + :param: int_depth: The depth estimate for the SML regression model + :param: int_scales: he scale to be used for the SML regression model + :param: device: The device on which the image would be allocated + :returns: The transformed tensor inputs suitable to be used with an SML model + """ + + sample = {'image' : input_image, 'int_depth' : int_depth, 'int_scales' : int_scales, 'int_depth_no_tf' : int_depth} + sample = scale_map_learner_transform(sample) + x = torch.cat([sample['int_depth'], sample['int_scales']], 0) + x = x.to(device) + d = sample['int_depth_no_tf'].to(device) + + return x.unsqueeze(0), d.unsqueeze(0) + +.. code:: ipython3 + + # Transform the dummy input image for the ScaleMapLearner model + # Note that this will lead to a tuple as an output. Both the elements + # which is fed to ScaleMapLearner during the conversion process to onxx + transformed_dummy_image_scale = transform_image_for_depth_scale(input_image=dummy_input, + scale_map_learner_transform=scale_map_learner_transform, + int_depth=d_depth, int_scales=d_scales) + +Conversion of Scale Map Learner model to OpenVINO™ IR format +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +The OpenVINO™ toolkit doesn’t provide any direct method of converting +PyTorch models to the intermediate representation format. To have the +associated ScaleMapLearner in the OpenVINO™ IR format and then compile +it, we shall follow the following steps: + +1. Load the model in memory via instantiating the + ``modules.midas.midas_net_custom.MidasNet_small_videpth`` class and + passing the downloaded checkpoint earlier as an argument. +2. Export the model to ``.onnx`` format using the transformed dummy + inputs created earlier. +3. Use the serialize function from OpenVINO to create equivalent + ``.xml`` and ``.bin`` files and obtain compiled models in the same + step. Alternatively serialization procedure may be avoided and + compiled model may be obtained by directly using OpenVINO’s + ``compile`` function. + +If the name of the ``.ckpt`` file is too much to handle, here is the +common format of all checkpoint files from the model releases. + +- sml_model.dpredictor..nsamples..ckpt +- Replace and with the depth estimation + model name and the no. of levels of depth density the SML model + has been trained on +- E.g. sml_model.dpredictor.dpt_hybrid.nsamples.500.ckpt will be the + file name corresponding to the SML model based on the dpt_hybrid + depth predictor and has been trained on 500 points of the density + level on the depth map + +.. code:: ipython3 + + # Run with the same min_pred and max_pred arguments which were used to compute + # global scale and shift alignment + scale_map_learner = MidasNet_small_videpth(path=str(MODEL_DIR / 'sml_model.dpredictor.midas_small.nsamples.1500.ckpt'), + min_pred=0.1, max_pred=8.0) + + +.. parsed-literal:: + + Loading weights: model/sml_model.dpredictor.midas_small.nsamples.1500.ckpt + + +.. parsed-literal:: + + Downloading: "https://github.com/rwightman/gen-efficientnet-pytorch/zipball/master" to model/master.zip + + +.. code:: ipython3 + + # As usual, since the MidasNet_small_videpthc class internally downloads a repo again from torch hub + # we shall clean the same since the model callable is now available to us + # Remove unnecessary directories and files and suppress errors(if any) + rmtree(path=str(MODEL_DIR / 'rwightman_gen-efficientnet-pytorch_master'), ignore_errors=True) + + # Check for the existence of the trusted list file and then remove + list_file = Path(MODEL_DIR / 'trusted_list') + if list_file.is_file(): + list_file.unlink() + +.. code:: ipython3 + + # Evaluate the model to switch some operations from training mode to inference. + scale_map_learner.eval() + + # Store the tuple of dummy variables into separate variables for easier reference + x_dummy, d_dummy = transformed_dummy_image_scale + + # Call the export function via the transformed dummy image obtained from the + # earlier steps. It is absolutely not a case of worry if multiple warnings pop up + # in this step. They can be safely ignored. + torch.onnx.export(model=scale_map_learner, args=(x_dummy, d_dummy), f=str(MODEL_DIR / 'scale_map_learner.onnx')) + + +.. parsed-literal:: + + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/246-depth-estimation-videpth/model/rwightman_gen-efficientnet-pytorch_master/geffnet/conv2d_layers.py:47: TracerWarning: Converting a tensor to a Python boolean might cause the trace to be incorrect. We can't record the data flow of Python values, so this value will be treated as a constant in the future. This means that the trace might not generalize to other inputs! + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/_internal/jit_utils.py:258: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + _C._jit_pass_onnx_node_shape_type_inference(node, params_dict, opset_version) + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) + _C._jit_pass_onnx_graph_shape_type_inference( + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:687: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + _C._jit_pass_onnx_graph_shape_type_inference( + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: Constant folding - Only steps=1 can be constant folded for opset >= 10 onnx::Slice op. Constant folding not applied. (Triggered internally at ../torch/csrc/jit/passes/onnx/constant_fold.cpp:179.) + _C._jit_pass_onnx_graph_shape_type_inference( + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torch/onnx/utils.py:1178: UserWarning: The shape inference of prim::Constant type is missing, so it may result in wrong shape inference for the exported graph. Please consider adding it in symbolic function. (Triggered internally at ../torch/csrc/jit/passes/onnx/shape_type_inference.cpp:1884.) + _C._jit_pass_onnx_graph_shape_type_inference( + + +Select inference device +''''''''''''''''''''''' + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +Compilation of the ScaleMapLearner(SML) model +''''''''''''''''''''''''''''''''''''''''''''' + +Now we can go ahead and compile our SML model from the ``.onnx`` file +path. We will not perform serialization because we don’t plan to re-read +the file again within this tutorial. Therefore we will use the compiled +SML model as it is. + +.. code:: ipython3 + + scale_map_learner = core.read_model(MODEL_DIR / 'scale_map_learner.onnx') + + # In the situation where you are unaware of the correct device to compile your + # model in, just set device_name='AUTO' and let OpenVINO decide for you + compiled_scale_map_learner = core.compile_model(model=scale_map_learner, device_name=device.value) + +.. code:: ipython3 + + def run_depth_scale_model(input_image_h: int, input_image_w: int, + transformed_image_for_depth_scale: Tuple[torch.Tensor, torch.Tensor], + compiled_scale_map_learner: type_compiled_model) -> np.ndarray: + + """ + Run the compiled_scale_map_learner on the transformed image of dimensions + input_image_w x input_image_h suitable to be used on such a model + + :param: input_image_h: The height of the input image + :param: input_image_w: The width of the input image + :param: transformed_image_for_depth_scale: The transformed image inputs suitable to be used as an input to the SML model + :returns: + sml_pred: The regression based prediction of the SML model + + """ + + # Obtain the last output layer separately + output_layer_scale_map_learner = compiled_scale_map_learner.output(0) + x_transform, d_transform = transformed_image_for_depth_scale + + with torch.no_grad(): + # Perform computation like a standard OpenVINO compiled model + sml_pred = torch.from_numpy(compiled_scale_map_learner([x_transform, d_transform])[output_layer_scale_map_learner]) + sml_pred = ( + torch.nn.functional.interpolate( + sml_pred, + size=(input_image_h, input_image_w), + mode='bicubic', + align_corners=False, + ) + .squeeze() + .cpu() + .numpy() + ) + + return sml_pred + +.. code:: ipython3 + + # Run the compiled SML model using the set of dummy inputs + sml_pred_dummy = run_depth_scale_model(input_image_h=IMAGE_H, input_image_w=IMAGE_W, + transformed_image_for_depth_scale=transformed_dummy_image_scale, + compiled_scale_map_learner=compiled_scale_map_learner) + +Storing and visualizing dummy results obtained +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +.. code:: ipython3 + + # Base directory in which outputs would be stored as a pathlib.Path variable + OUTPUT_DIR = Path('output') + + # Create the output directory adjacent to the notebook and suppress errors if the directory already exists + OUTPUT_DIR.mkdir(exist_ok=True) + + # Utility functions are directly available in modules.midas.utils + # Provide path names without any extension and let the write_depth + # function provide them for you. Take note of the arguments. + utils.write_depth(path=str(OUTPUT_DIR / 'dummy_input'), depth=d_depth, bits=2) + utils.write_depth(path=str(OUTPUT_DIR / 'dummy_input_sml'), depth=sml_pred_dummy, bits=2) + +.. code:: ipython3 + + plt.figure() + + img_dummy_in = mpimg.imread('data/image/dummy_img.png') + img_dummy_out = mpimg.imread(OUTPUT_DIR / 'dummy_input.png') + img_dummy_sml_out = mpimg.imread(OUTPUT_DIR / 'dummy_input_sml.png') + + f, axes = plt.subplots(1, 3) + plt.subplots_adjust(right=2.0) + + axes[0].imshow(img_dummy_in) + axes[1].imshow(img_dummy_out) + axes[2].imshow(img_dummy_sml_out) + + axes[0].set_title('dummy input') + axes[1].set_title('depth prediction on dummy input') + axes[2].set_title('SML on depth estimate') + + + + +.. parsed-literal:: + + Text(0.5, 1.0, 'SML on depth estimate') + + + + +.. parsed-literal:: + +
+ + + +.. image:: 246-depth-estimation-videpth-with-output_files/246-depth-estimation-videpth-with-output_48_2.png + + +Running inference on a test image +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +Now role of both the dummy inputs i.e. the dummy image as well as its +associated depth map is now over. Since we have access to the compiled +models now, we can load the *one* image available to us for pure +inferencing purposes and run all the above steps one by one till +plotting of the depth map. + +If you haven’t noticed already the data directory of this tutorial has +been arranged as follows. This allows us to comply to these +`rules `__. + +.. code:: bash + + data + ├── image + │ ├── dummy_img.png # RGB images + │ └── .png + └── sparse_depth + ├── dummy_img.png # sparse metric depth maps + └── .png # as 16b PNG files + +At the same time, the depth storage method `used in the VOID +dataset `__ +is assumed. + +If you are thinking of the file name format of the image for inference, +here is the reasoning. + +The dataset was collected using the Intel `RealSense D435i +camera `__, which was +configured to produce synchronized accelerometer and gyroscope +measurements at 400 Hz, along with synchronized VGA-size (640 x 480) RGB +and depth streams at 30 Hz. The depth frames are acquired using active +stereo and is aligned to the RGB frame using the sensor factory +calibration. The frequency of sensor and depth stream input run at +certain fixed frequencies and hence time stamping every frame captured +is beneficial for maintaining structure as well as for debugging +purposes later. + +*The image for inference and it sparse depth map is taken from the +compressed dataset +present*\ `here `__ + +.. code:: ipython3 + + # As before download the sample images for inference and take note of the image hashes if you + # want to use them later + download_file('https://user-images.githubusercontent.com/22426058/254174393-fc6dcc5f-f677-4618-b2ef-22e8e5cb1ebe.png', filename='1552097950.2672.png', directory=Path(DATA_DIR / 'image'), silent=True) + download_file('https://user-images.githubusercontent.com/22426058/254174379-5d00b66b-57b4-4e96-91e9-36ef15ec5a0a.png', filename='1552097950.2672.png', directory=Path(DATA_DIR / 'sparse_depth'), silent=True) + + # Load the image and its depth scale + img_input = data_loader.load_input_image('data/image/1552097950.2672.png') + img_depth_input = data_loader.load_sparse_depth('data/sparse_depth/1552097950.2672.png') + + # Transform the input image for the depth model + transformed_image = transform_image_for_depth(input_image=img_input, depth_model_transform=depth_model_transform) + + # Run the depth model on the transformed input + depth_pred = run_depth_model(input_image_h=IMAGE_H, input_image_w=IMAGE_W, + transformed_image=transformed_image, compiled_depth_model=compiled_depth_model) + + + # Call the function on the sparse depth map + # with all default settings and store in appropriate variables + int_depth, int_scales = compute_global_scale_and_shift(input_sparse_depth=img_depth_input, validity_map=None, depth_pred=depth_pred) + + # Transform the input image for the ScaleMapLearner model + transformed_image_scale = transform_image_for_depth_scale(input_image=img_input, + scale_map_learner_transform=scale_map_learner_transform, + int_depth=int_depth, int_scales=int_scales) + + # Run the SML model using the set of inputs + sml_pred = run_depth_scale_model(input_image_h=IMAGE_H, input_image_w=IMAGE_W, + transformed_image_for_depth_scale=transformed_image_scale, + compiled_scale_map_learner=compiled_scale_map_learner) + + + +.. parsed-literal:: + + data/image/1552097950.2672.png: 0%| | 0.00/371k [00:00 + + + +.. image:: 246-depth-estimation-videpth-with-output_files/246-depth-estimation-videpth-with-output_53_2.png + + +Cleaning up the data directory +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +We will *follow suit* for the directory in which we downloaded images +and depth maps from another repo. We shall move remove the unnecessary +directories and files which were created during the download process. + +.. code:: ipython3 + + # Remove the data directory and suppress errors(if any) + rmtree(path=str(DATA_DIR), ignore_errors=True) + +Concluding notes +~~~~~~~~~~~~~~~~ + + 1. The code for this tutorial is adapted from the `VI-Depth + repository `__. + 2. Users may choose to download the original and raw datasets from + the `VOID + dataset `__. + 3. The `isl-org/VI-Depth `__ + works on a slightly older version of released model assets from + its `MiDaS sibling + repository `__. However, the new + releases beginning from + `v3.1 `__ + directly have OpenVINO™ ``.xml`` and ``.bin`` model files as their + assets thereby rendering the **major pre-processing and model + compilation step irrelevant**. diff --git a/docs/notebooks/246-depth-estimation-videpth-with-output_files/246-depth-estimation-videpth-with-output_48_2.png b/docs/notebooks/246-depth-estimation-videpth-with-output_files/246-depth-estimation-videpth-with-output_48_2.png new file mode 100644 index 00000000000..b21ee5ba8cc --- /dev/null +++ b/docs/notebooks/246-depth-estimation-videpth-with-output_files/246-depth-estimation-videpth-with-output_48_2.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:156808b0fbe49b48f340663ec6bf41222574b3a6afd0657473a9e6a3c3a3fb7e +size 215788 diff --git a/docs/notebooks/246-depth-estimation-videpth-with-output_files/246-depth-estimation-videpth-with-output_53_2.png b/docs/notebooks/246-depth-estimation-videpth-with-output_files/246-depth-estimation-videpth-with-output_53_2.png new file mode 100644 index 00000000000..90ddeee778d --- /dev/null +++ b/docs/notebooks/246-depth-estimation-videpth-with-output_files/246-depth-estimation-videpth-with-output_53_2.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:f552d67a7af53d7d7f4175d00206aebff176df393889b1543a5b9c6bd2938894 +size 190117 diff --git a/docs/notebooks/246-depth-estimation-videpth-with-output_files/index.html b/docs/notebooks/246-depth-estimation-videpth-with-output_files/index.html new file mode 100644 index 00000000000..2e55e6f38e2 --- /dev/null +++ b/docs/notebooks/246-depth-estimation-videpth-with-output_files/index.html @@ -0,0 +1,8 @@ + +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/246-depth-estimation-videpth-with-output_files/ + +

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/246-depth-estimation-videpth-with-output_files/


../
+246-depth-estimation-videpth-with-output_48_2.png  16-Aug-2023 01:31              215788
+246-depth-estimation-videpth-with-output_53_2.png  16-Aug-2023 01:31              190117
+

+ diff --git a/docs/notebooks/247-code-language-id-with-output.rst b/docs/notebooks/247-code-language-id-with-output.rst new file mode 100644 index 00000000000..68fb8ba5b90 --- /dev/null +++ b/docs/notebooks/247-code-language-id-with-output.rst @@ -0,0 +1,683 @@ +Programming Language Classification with OpenVINO +================================================= + +Overview +-------- + +This tutorial will be divided in 2 parts: 1. Create a simple inference +pipeline with a pre-trained model using the OpenVINO™ IR format. 2. +Conduct `post-training quantization `__ +on a pre-trained model using Hugging Face Optimum and benchmark +performance. + +Feel free to use the notebook outline in Jupyter or your IDE for easy +navigation. + +Introduction +------------ + +Task +~~~~ + +**Programming language classification** is the task of identifying which +programming language is used in an arbitrary code snippet. This can be +useful to label new data to include in a dataset, and potentially serve +as an intermediary step when input snippets need to be process based on +their programming language. + +It is a relatively easy machine learning task given that each +programming language has its own formal symbols, syntax, and grammar. +However, there are some potential edge cases: - **Ambiguous short +snippets**: For example, TypeScript is a superset of JavaScript, meaning +it does everything JavaScript can and more. For a short input snippet, +it might be impossible to distinguish between the two. Given we know +TypeScript is a superset, and the model doesn’t, we should default to +classifying the input as JavaScript in a post-processing step. - +**Nested programming languages**: Some languages are typically used in +tandem. For example, most HTML contains CSS and JavaScript, and it is +not uncommon to see SQL nested in other scripting languages. For such +input, it is unclear what the expected output class should be. - +**Evolving programming language**: Even though programming languages are +formal, their symbols, syntax, and grammar can be revised and updated. +For example, the walrus operator (``:=``) was a symbol distinctively +used in Golang, but was later introduced in Python 3.8. + +Model +~~~~~ + +The classification model that will be used in this notebook is +`CodeBERTa-language-id `__ +by HuggingFace. This model was fine-tuned from the masked language +modeling model +`CodeBERTa-small-v1 `__ +trained on the +`CodeSearchNet `__ +dataset (Husain, 2019). + +It supports 6 programming languages: - Go - Java - JavaScript - PHP - +Python - Ruby + +Part 1: Inference pipeline with OpenVINO +---------------------------------------- + +For this section, we will use the `HuggingFace Optimum `__ library, which +aims to optimize inference on specific hardware and integrates with the +OpenVINO toolkit. The code will be very similar to the +`HuggingFace Transformers `__, but +will allow to automatically convert models to the OpenVINO™ IR format. + +Install prerequisites +~~~~~~~~~~~~~~~~~~~~~ + +First, complete the `repository installation steps <../../README.md>`__. + +Then, the following cell will install: - HuggingFace Optimum with +OpenVINO support - HuggingFace Evaluate to benchmark results + +.. code:: ipython3 + + !pip install -q "diffusers>=0.17.1" "openvino-dev>=2023.0.0" "nncf>=2.5.0" "gradio" "onnx>=1.11.0" "onnxruntime>=1.14.0" "optimum-intel>=1.9.1" "transformers>=4.31.0" "evaluate" + + +.. parsed-literal:: + + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + ERROR: pip's dependency resolver does not currently take into account all the packages that are installed. This behaviour is the source of the following dependency conflicts. + pytorch-lightning 1.6.5 requires protobuf<=3.20.1, but you have protobuf 4.24.0 which is incompatible. + + +Imports +~~~~~~~ + +The import ``OVModelForSequenceClassification`` from Optimum is +equivalent to ``AutoModelForSequenceClassification`` from Transformers + +.. code:: ipython3 + + from functools import partial + from pathlib import Path + + import pandas as pd + from datasets import load_dataset, Dataset + import evaluate + from transformers import pipeline, AutoTokenizer, AutoModelForSequenceClassification + from optimum.intel import OVModelForSequenceClassification + from optimum.intel.openvino import OVConfig, OVQuantizer + from huggingface_hub.utils import RepositoryNotFoundError + + +.. parsed-literal:: + + 2023-08-16 01:03:40.095980: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-08-16 01:03:40.129769: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. + 2023-08-16 01:03:40.709247: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + + +.. parsed-literal:: + + INFO:nncf:NNCF initialized successfully. Supported frameworks detected: torch, tensorflow, onnx, openvino + + +.. parsed-literal:: + + No CUDA runtime is found, using CUDA_HOME='/usr/local/cuda' + + +Setting up HuggingFace cache +~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +Resources from HuggingFace will be downloaded in the local folder +``./model`` (next to this notebook) instead of the device global cache +for easy cleanup. Learn more +`here `__. + +.. code:: ipython3 + + MODEL_NAME = "CodeBERTa-language-id" + MODEL_ID = f"huggingface/{MODEL_NAME}" + MODEL_LOCAL_PATH = Path("./model").joinpath(MODEL_NAME) + +Select inference device +~~~~~~~~~~~~~~~~~~~~~~~ + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + from openvino.runtime import Core + + core = Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +Download resources +~~~~~~~~~~~~~~~~~~ + +.. code:: ipython3 + + # try to load resources locally + try: + model = OVModelForSequenceClassification.from_pretrained(MODEL_LOCAL_PATH, device=device.value) + tokenizer = AutoTokenizer.from_pretrained(MODEL_LOCAL_PATH) + print(f"Loaded resources from local path: {MODEL_LOCAL_PATH.absolute()}") + + # if not found, download from HuggingFace Hub then save locally + except (RepositoryNotFoundError, OSError): + print("Downloading resources from HuggingFace Hub") + tokenizer = AutoTokenizer.from_pretrained(MODEL_ID) + tokenizer.save_pretrained(MODEL_LOCAL_PATH) + + # export=True is needed to convert the PyTorch model to OpenVINO + model = OVModelForSequenceClassification.from_pretrained(MODEL_ID, export=True, device=device.value) + model.save_pretrained(MODEL_LOCAL_PATH) + print(f"Ressources cached locally at: {MODEL_LOCAL_PATH.absolute()}") + + +.. parsed-literal:: + + Downloading resources from HuggingFace Hub + + +.. parsed-literal:: + + Framework not specified. Using pt to export to ONNX. + Some weights of the model checkpoint at huggingface/CodeBERTa-language-id were not used when initializing RobertaForSequenceClassification: ['roberta.pooler.dense.bias', 'roberta.pooler.dense.weight'] + - This IS expected if you are initializing RobertaForSequenceClassification from the checkpoint of a model trained on another task or with another architecture (e.g. initializing a BertForSequenceClassification model from a BertForPreTraining model). + - This IS NOT expected if you are initializing RobertaForSequenceClassification from the checkpoint of a model that you expect to be exactly identical (initializing a BertForSequenceClassification model from a BertForSequenceClassification model). + Using framework PyTorch: 1.13.1+cpu + Overriding 1 configuration item(s) + - use_cache -> False + Compiling the model... + Set CACHE_DIR to /tmp/tmpsl_db7y_/model_cache + + +.. parsed-literal:: + + Ressources cached locally at: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/247-code-language-id/model/CodeBERTa-language-id + + +Create inference pipeline +~~~~~~~~~~~~~~~~~~~~~~~~~ + +.. code:: ipython3 + + code_classification_pipe = pipeline("text-classification", model=model, tokenizer=tokenizer) + + +.. parsed-literal:: + + Xformers is not installed correctly. If you want to use memory_efficient_attention to accelerate training use the following command to install Xformers + pip install xformers. + + +Inference on new input +~~~~~~~~~~~~~~~~~~~~~~ + +.. code:: ipython3 + + # change input snippet to test model + input_snippet = "df['speed'] = df.distance / df.time" + output = code_classification_pipe(input_snippet) + + print(f"Input snippet:\n {input_snippet}\n") + print(f"Predicted label: {output[0]['label']}") + print(f"Predicted score: {output[0]['score']:.2}") + + +.. parsed-literal:: + + Input snippet: + df['speed'] = df.distance / df.time + + Predicted label: python + Predicted score: 0.81 + + +Part 2: OpenVINO post-training quantization with HuggingFace Optimum +-------------------------------------------------------------------- + +In this section, we will quantize a trained model. At a high-level, this +process consists of using lower precision numbers in the model, which +results in a smaller model size and faster inference at the cost of a +potential marginal performance degradation. `Learn more `__. + +The HuggingFace Optimum library supports post-training quantization for +OpenVINO. `Learn more `__. + +Define constants and functions +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +.. code:: ipython3 + + QUANTIZED_MODEL_LOCAL_PATH = MODEL_LOCAL_PATH.with_name(f"{MODEL_NAME}-quantized") + DATASET_NAME = "code_search_net" + LABEL_MAPPING = {"go": 0, "java": 1, "javascript": 2, "php": 3, "python": 4, "ruby": 5} + + + def preprocess_function(examples: dict, tokenizer): + """Preprocess inputs by tokenizing the `func_code_string` column""" + return tokenizer( + examples["func_code_string"], + padding="max_length", + max_length=tokenizer.model_max_length, + truncation=True, + ) + + + def map_labels(example: dict) -> dict: + """Convert string labels to integers""" + label_mapping = {"go": 0, "java": 1, "javascript": 2, "php": 3, "python": 4, "ruby": 5} + example["language"] = label_mapping[example["language"]] + return example + + + def get_dataset_sample(dataset_split: str, num_samples: int) -> Dataset: + """Create a sample with equal representation of each class without downloading the entire data""" + labels = ["go", "java", "javascript", "php", "python", "ruby"] + example_per_label = num_samples // len(labels) + + examples = [] + for label in labels: + subset = load_dataset("code_search_net", split=dataset_split, name=label, streaming=True) + subset = subset.map(map_labels) + examples.extend([example for example in subset.shuffle().take(example_per_label)]) + + return Dataset.from_list(examples) + +Load resources +~~~~~~~~~~~~~~ + +NOTE: the base model is loaded using +``AutoModelForSequenceClassification`` from ``Transformers`` + +.. code:: ipython3 + + tokenizer = AutoTokenizer.from_pretrained(MODEL_LOCAL_PATH) + base_model = AutoModelForSequenceClassification.from_pretrained(MODEL_ID) + + quantizer = OVQuantizer.from_pretrained(base_model) + quantization_config = OVConfig() + + +.. parsed-literal:: + + Some weights of the model checkpoint at huggingface/CodeBERTa-language-id were not used when initializing RobertaForSequenceClassification: ['roberta.pooler.dense.bias', 'roberta.pooler.dense.weight'] + - This IS expected if you are initializing RobertaForSequenceClassification from the checkpoint of a model trained on another task or with another architecture (e.g. initializing a BertForSequenceClassification model from a BertForPreTraining model). + - This IS NOT expected if you are initializing RobertaForSequenceClassification from the checkpoint of a model that you expect to be exactly identical (initializing a BertForSequenceClassification model from a BertForSequenceClassification model). + + +Load calibration dataset +~~~~~~~~~~~~~~~~~~~~~~~~ + +The ``get_dataset_sample()`` function will sample up to ``num_samples``, +with an equal number of examples across the 6 programming languages. + +NOTE: Uncomment the method below to download and use the full dataset +(5+ Gb). + +.. code:: ipython3 + + calibration_sample = get_dataset_sample(dataset_split="train", num_samples=120) + calibration_sample = calibration_sample.map(partial(preprocess_function, tokenizer=tokenizer)) + + # calibration_sample = quantizer.get_calibration_dataset( + # DATASET_NAME, + # preprocess_function=partial(preprocess_function, tokenizer=tokenizer), + # num_samples=120, + # dataset_split="train", + # preprocess_batch=True, + # ) + + +.. parsed-literal:: + + huggingface/tokenizers: The current process just got forked, after parallelism has already been used. Disabling parallelism to avoid deadlocks... + To disable this warning, you can either: + - Avoid using `tokenizers` before the fork if possible + - Explicitly set the environment variable TOKENIZERS_PARALLELISM=(true | false) + + + +.. parsed-literal:: + + Map: 0%| | 0/120 [00:00 False + + +.. parsed-literal:: + + WARNING:nncf:You are setting `forward` on an NNCF-processed model object. + NNCF relies on custom-wrapping the `forward` call in order to function properly. + Arbitrary adjustments to the forward function on an NNCFNetwork object have undefined behaviour. + If you need to replace the underlying forward function of the original model so that NNCF should be using that instead of the original forward function that NNCF saved during the compressed model creation, you can do this by calling: + model.nncf.set_original_unbound_forward(fn) + if `fn` has an unbound 0-th `self` argument, or + with model.nncf.temporary_bound_original_forward(fn): ... + if `fn` already had 0-th `self` argument bound or never had it in the first place. + WARNING:nncf:You are setting `forward` on an NNCF-processed model object. + NNCF relies on custom-wrapping the `forward` call in order to function properly. + Arbitrary adjustments to the forward function on an NNCFNetwork object have undefined behaviour. + If you need to replace the underlying forward function of the original model so that NNCF should be using that instead of the original forward function that NNCF saved during the compressed model creation, you can do this by calling: + model.nncf.set_original_unbound_forward(fn) + if `fn` has an unbound 0-th `self` argument, or + with model.nncf.temporary_bound_original_forward(fn): ... + if `fn` already had 0-th `self` argument bound or never had it in the first place. + + +.. parsed-literal:: + + Configuration saved in model/CodeBERTa-language-id-quantized/openvino_config.json + + +Load quantized model +~~~~~~~~~~~~~~~~~~~~ + +NOTE: the argument ``export=True`` is not required since the quantized +model is already in the OpenVINO format. + +.. code:: ipython3 + + quantized_model = OVModelForSequenceClassification.from_pretrained(QUANTIZED_MODEL_LOCAL_PATH, device=device.value) + quantized_code_classification_pipe = pipeline("text-classification", model=quantized_model, tokenizer=tokenizer) + + +.. parsed-literal:: + + Compiling the model... + Set CACHE_DIR to model/CodeBERTa-language-id-quantized/model_cache + + +Inference on new input using quantized model +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +.. code:: ipython3 + + input_snippet = "df['speed'] = df.distance / df.time" + output = quantized_code_classification_pipe(input_snippet) + + print(f"Input snippet:\n {input_snippet}\n") + print(f"Predicted label: {output[0]['label']}") + print(f"Predicted score: {output[0]['score']:.2}") + + +.. parsed-literal:: + + Input snippet: + df['speed'] = df.distance / df.time + + Predicted label: python + Predicted score: 0.82 + + +Load evaluation set +~~~~~~~~~~~~~~~~~~~ + +NOTE: Uncomment the method below to download and use the full dataset +(5+ Gb). + +.. code:: ipython3 + + validation_sample = get_dataset_sample(dataset_split="validation", num_samples=120) + + # validation_sample = load_dataset(DATASET_NAME, split="validation") + +Evaluate model +~~~~~~~~~~~~~~ + +.. code:: ipython3 + + # This class is needed due to a current limitation of the Evaluate library with multiclass metrics + # ref: https://discuss.huggingface.co/t/combining-metrics-for-multiclass-predictions-evaluations/21792/16 + class ConfiguredMetric: + def __init__(self, metric, *metric_args, **metric_kwargs): + self.metric = metric + self.metric_args = metric_args + self.metric_kwargs = metric_kwargs + + def add(self, *args, **kwargs): + return self.metric.add(*args, **kwargs) + + def add_batch(self, *args, **kwargs): + return self.metric.add_batch(*args, **kwargs) + + def compute(self, *args, **kwargs): + return self.metric.compute(*args, *self.metric_args, **kwargs, **self.metric_kwargs) + + @property + def name(self): + return self.metric.name + + def _feature_names(self): + return self.metric._feature_names() + +First, an ``Evaluator`` object for ``text-classification`` and a set of +``EvaluationModule`` are instantiated. Then, the evaluator +``.compute()`` method is called on both the base +``code_classification_pipe`` and the quantized +``quantized_code_classification_pipeline``. Finally, results are +displayed. + +.. code:: ipython3 + + code_classification_evaluator = evaluate.evaluator("text-classification") + # instantiate an object that can contain multiple `evaluate` metrics + metrics = evaluate.combine([ + ConfiguredMetric(evaluate.load('f1'), average='macro'), + ]) + + base_results = code_classification_evaluator.compute( + model_or_pipeline=code_classification_pipe, + data=validation_sample, + input_column="func_code_string", + label_column="language", + label_mapping=LABEL_MAPPING, + metric=metrics, + ) + + quantized_results = code_classification_evaluator.compute( + model_or_pipeline=quantized_code_classification_pipe, + data=validation_sample, + input_column="func_code_string", + label_column="language", + label_mapping=LABEL_MAPPING, + metric=metrics, + ) + + results_df = pd.DataFrame.from_records([base_results, quantized_results], index=["base", "quantized"]) + results_df + + + + +.. raw:: html + +
+ + + + + + + + + + + + + + + + + + + + + + + + + + + +
f1total_time_in_secondssamples_per_secondlatency_in_seconds
base1.02.32259351.6663920.019355
quantized1.02.64746645.3263570.022062
+
+ + + +Additional resources +-------------------- + +- `Grammatical Error Correction with + OpenVINO `__ +- `Quantize a Hugging Face Question-Answering Model with + OpenVINO `__\ \*\* + +Clean up +-------- + +Uncomment and run cell below to delete all resources cached locally in +./model + +.. code:: ipython3 + + # import os + # import shutil + + # try: + # shutil.rmtree(path=QUANTIZED_MODEL_LOCAL_PATH) + # shutil.rmtree(path=MODEL_LOCAL_PATH) + # os.remove(path="./compressed_graph.dot") + # os.remove(path="./original_graph.dot") + # except FileNotFoundError: + # print("Directory was already deleted") diff --git a/docs/notebooks/248-stable-diffusion-xl-with-output.rst b/docs/notebooks/248-stable-diffusion-xl-with-output.rst new file mode 100644 index 00000000000..5631ca77c0c --- /dev/null +++ b/docs/notebooks/248-stable-diffusion-xl-with-output.rst @@ -0,0 +1,612 @@ +Image generation with Stable Diffusion XL and OpenVINO +====================================================== + +.. _top: + +Stable Diffusion XL or SDXL is the latest image generation model that is +tailored towards more photorealistic outputs with more detailed imagery +and composition compared to previous Stable Diffusion models, including +Stable Diffusion 2.1. + +With Stable Diffusion XL you can now make more realistic images with +improved face generation, produce legible text within images, and create +more aesthetically pleasing art using shorter prompts. + +.. figure:: https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0/resolve/main/pipeline.png + :alt: pipeline + + pipeline + +`SDXL `__ consists of an `ensemble of +experts `__ pipeline for latent +diffusion: In the first step, the base model is used to generate (noisy) +latents, which are then further processed with a refinement model +specialized for the final denoising steps. Note that the base model can +be used as a standalone module or in a two-stage pipeline as follows: +First, the base model is used to generate latents of the desired output +size. In the second step, we use a specialized high-resolution model and +apply a technique called +`SDEdit `__\ ( also known as “image to +image”) to the latents generated in the first step, using the same +prompt. + +Compared to previous versions of Stable Diffusion, SDXL leverages a +three times larger UNet backbone: The increase of model parameters is +mainly due to more attention blocks and a larger cross-attention context +as SDXL uses a second text encoder. The authors design multiple novel +conditioning schemes and train SDXL on multiple aspect ratios and also +introduce a refinement model that is used to improve the visual fidelity +of samples generated by SDXL using a post-hoc image-to-image technique. +The testing of SDXL shows drastically improved performance compared to +the previous versions of Stable Diffusion and achieves results +competitive with those of black-box state-of-the-art image generators. + +In this tutorial, we consider how to run the SDXL model using OpenVINO. + +We will use a pre-trained model from the `Hugging Face +Diffusers `__ library. To +simplify the user experience, the `Hugging Face Optimum +Intel `__ library is +used to convert the models to OpenVINO™ IR format. + +The tutorial consists of the following steps: + +- Install prerequisites +- Download the Stable Diffusion XL Base model from a public source + using the `OpenVINO integration with Hugging Face + Optimum `__. +- Run Text2Image generation pipeline using Stable Diffusion XL base +- Run Image2Image generation pipeline using Stable Diffusion XL base +- Download and convert the Stable Diffusion XL Refiner model from a + public source using the `OpenVINO integration with Hugging Face + Optimum `__. +- Run 2-stages Stable Diffusion XL pipeline + +.. + + **Note**: Some demonstrated models can require at least 64GB RAM for + conversion and running. + +**Table of contents**: + +- `Install Prerequisites <#install-prerequisites>`__ +- `SDXL Base model <#sdxl-base-model>`__ + + - `Select inference device <#select-inference-device>`__ + - `Run Text2Image generation pipeline <#run-text2image-generation-pipeline>`__ + - `Text2image Generation Interactive Demo <#text2image-generation-interactive-demo>`__ + - `Run Image2Image generation pipeline <#run-image2image-generation-pipeline>`__ + - `Image2Image Generation Interactive Demo <#image2image-generation-interactive-demo>`__ + +- `SDXL Refiner model <#sdxl-refiner-model>`__ + + - `Select inference device <#select-inference-device>`__ + - `Run Text2Image generation with Refinement <#run-text2image-generation-with-refinement>`__ + +Install prerequisites\ `⇑ <#top>`__ +############################################################################################################################### + + +.. code:: ipython3 + + !pip install -q "git+https://github.com/huggingface/optimum-intel.git" + !pip install -q "openvino-dev==2023.1.0.dev20230728" + !pip install -q --upgrade-strategy eager "diffusers>=0.18.0" "invisible-watermark>=0.2.0" "transformers>=4.30.2" "accelerate" "onnx" "onnxruntime" + !pip install -q gradio + +SDXL Base model\ `⇑ <#top>`__ +############################################################################################################################### + + +We will start with the base model part, which is responsible for the +generation of images of the desired output size. +`stable-diffusion-xl-base-1.0 `__ +is available for downloading via the `HuggingFace +hub `__. It already provides a +ready-to-use model in OpenVINO format compatible with `Optimum +Intel `__. + +To load an OpenVINO model and run an inference with OpenVINO Runtime, +you need to replace diffusers ``StableDiffusionXLPipeline`` with Optimum +``OVStableDiffusionXLPipeline``. In case you want to load a PyTorch +model and convert it to the OpenVINO format on the fly, you can set +``export=True``. + +You can save the model on disk using the ``save_pretrained`` method. + +.. code:: ipython3 + + from pathlib import Path + from optimum.intel.openvino import OVStableDiffusionXLPipeline + import gc + + model_id = "stabilityai/stable-diffusion-xl-base-1.0" + model_dir = Path("openvino-sd-xl-base-1.0") + + +.. parsed-literal:: + + 2023-08-06 20:21:05.073866: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-08-06 20:21:05.114013: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. + 2023-08-06 20:21:05.843627: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + + +.. parsed-literal:: + + INFO:nncf:NNCF initialized successfully. Supported frameworks detected: torch, tensorflow, onnx, openvino + + +.. parsed-literal:: + + No CUDA runtime is found, using CUDA_HOME='/usr/local/cuda' + + +Select inference device\ `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + from openvino.runtime import Core + + core = Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=2, options=('CPU', 'GPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + if not model_dir.exists(): + text2image_pipe = OVStableDiffusionXLPipeline.from_pretrained(model_id, compile=False, device=device.value) + text2image_pipe.half() + text2image_pipe.save_pretrained(model_dir) + text2image_pipe.compile() + else: + text2image_pipe = OVStableDiffusionXLPipeline.from_pretrained(model_dir, device=device.value) + + +.. parsed-literal:: + + Compiling the vae_decoder... + Compiling the unet... + Compiling the text_encoder_2... + Compiling the text_encoder... + Compiling the vae_encoder... + + +Run Text2Image generation pipeline\ `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Now, we can run the model for the generation of images using text +prompts. To speed up evaluation and reduce the required memory we +decrease ``num_inference_steps`` and image size (using ``height`` and +``width``). You can modify them to suit your needs and depend on the +target hardware. We also specified a ``generator`` parameter based on a +numpy random state with a specific seed for results reproducibility. + +.. code:: ipython3 + + import numpy as np + + prompt = "cute cat 4k, high-res, masterpiece, best quality, soft lighting, dynamic angle" + image = text2image_pipe(prompt, num_inference_steps=15, height=512, width=512, generator=np.random.RandomState(314)).images[0] + image.save("cat.png") + image + + +.. parsed-literal:: + + /home/ea/work/ov_notebooks_env/lib/python3.8/site-packages/optimum/intel/openvino/modeling_diffusion.py:552: FutureWarning: `shared_memory` is deprecated and will be removed in 2024.0. Value of `shared_memory` is going to override `share_inputs` value. Please use only `share_inputs` explicitly. + outputs = self.request(inputs, shared_memory=True) + + + +.. parsed-literal:: + + 0%| | 0/15 [00:00`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +.. code:: ipython3 + + import gradio as gr + + if text2image_pipe is None: + text2image_pipe = OVStableDiffusionXLPipeline.from_pretrained(model_dir, device=device.value) + + prompt = "cute cat 4k, high-res, masterpiece, best quality, soft lighting, dynamic angle" + + def generate_from_text(text, seed, num_steps): + result = text2image_pipe(text, num_inference_steps=num_steps, generator=np.random.RandomState(seed), height=512, width=512).images[0] + return result + + + with gr.Blocks() as demo: + with gr.Column(): + positive_input = gr.Textbox(label="Text prompt") + with gr.Row(): + seed_input = gr.Number(precision=0, label="Seed", value=42, minimum=0) + steps_input = gr.Slider(label="Steps", value=10) + btn = gr.Button() + out = gr.Image(label="Result", type="pil", width=512) + btn.click(generate_from_text, [positive_input, seed_input, steps_input], out) + gr.Examples([ + [prompt, 999, 20], + ["underwater world coral reef, colorful jellyfish, 35mm, cinematic lighting, shallow depth of field, ultra quality, masterpiece, realistic", 89, 20], + ["a photo realistic happy white poodle dog ​​playing in the grass, extremely detailed, high res, 8k, masterpiece, dynamic angle", 1569, 15], + ["Astronaut on Mars watching sunset, best quality, cinematic effects,", 65245, 12], + ["Black and white street photography of a rainy night in New York, reflections on wet pavement", 48199, 10] + ], [positive_input, seed_input, steps_input]) + + # if you are launching remotely, specify server_name and server_port + # demo.launch(server_name='your server name', server_port='server port in int') + # Read more in the docs: https://gradio.app/docs/ + # if you want create public link for sharing demo, please add share=True + demo.launch() + + +.. parsed-literal:: + + Running on local URL: http://127.0.0.1:7860 + + To create a public link, set `share=True` in `launch()`. + + + +.. raw:: html + +
+ + +.. code:: ipython3 + + demo.close() + text2image_pipe = None + gc.collect(); + + +.. parsed-literal:: + + Closing server running on port: 7860 + + +Run Image2Image generation pipeline\ `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +We can reuse the already converted model for running the Image2Image +generation pipeline. For that, we should replace +``OVStableDiffusionXLPipeline`` with +``OVStableDiffusionXLImage2ImagePipeline``. + +Select inference device +^^^^^^^^^^^^^^^^^^^^^^^ + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=2, options=('CPU', 'GPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + from optimum.intel import OVStableDiffusionXLImg2ImgPipeline + + image2image_pipe = OVStableDiffusionXLImg2ImgPipeline.from_pretrained(model_dir, device=device.value) + + +.. parsed-literal:: + + Compiling the vae_decoder... + Compiling the unet... + Compiling the text_encoder... + Compiling the vae_encoder... + Compiling the text_encoder_2... + + +.. code:: ipython3 + + photo_prompt = "professional photo of a cat, extremely detailed, hyper realistic, best quality, full hd" + photo_image = image2image_pipe(photo_prompt, image=image, num_inference_steps=25, generator=np.random.RandomState(356)).images[0] + photo_image.save("photo_cat.png") + photo_image + + +.. parsed-literal:: + + /home/ea/work/ov_notebooks_env/lib/python3.8/site-packages/optimum/intel/openvino/modeling_diffusion.py:552: FutureWarning: `shared_memory` is deprecated and will be removed in 2024.0. Value of `shared_memory` is going to override `share_inputs` value. Please use only `share_inputs` explicitly. + outputs = self.request(inputs, shared_memory=True) + /home/ea/work/ov_notebooks_env/lib/python3.8/site-packages/optimum/pipelines/diffusers/pipeline_utils.py:64: FutureWarning: The preprocess method is deprecated and will be removed in a future version. Please use VaeImageProcessor.preprocess instead + warnings.warn( + /home/ea/work/ov_notebooks_env/lib/python3.8/site-packages/optimum/intel/openvino/modeling_diffusion.py:615: FutureWarning: `shared_memory` is deprecated and will be removed in 2024.0. Value of `shared_memory` is going to override `share_inputs` value. Please use only `share_inputs` explicitly. + outputs = self.request(inputs, shared_memory=True) + + + +.. parsed-literal:: + + 0%| | 0/7 [00:00`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +.. code:: ipython3 + + import gradio as gr + from diffusers.utils import load_image + import numpy as np + + + load_image( + "https://huggingface.co/datasets/optimum/documentation-images/resolve/main/intel/openvino/sd_xl/castle_friedrich.png" + ).resize((512, 512)).save("castle_friedrich.png") + + + if image2image_pipe is None: + image2image_pipe = OVStableDiffusionXLImg2ImgPipeline.from_pretrained(model_dir) + + def generate_from_image(text, image, seed, num_steps): + result = image2image_pipe(text, image=image, num_inference_steps=num_steps, generator=np.random.RandomState(seed)).images[0] + return result + + + with gr.Blocks() as demo: + with gr.Column(): + positive_input = gr.Textbox(label="Text prompt") + with gr.Row(): + seed_input = gr.Number(precision=0, label="Seed", value=42, minimum=0) + steps_input = gr.Slider(label="Steps", value=10) + btn = gr.Button() + with gr.Row(): + i2i_input = gr.Image(label="Input image", type="pil") + out = gr.Image(label="Result", type="pil", width=512) + btn.click(generate_from_image, [positive_input, i2i_input, seed_input, steps_input], out) + gr.Examples([ + ["amazing landscape from legends", "castle_friedrich.png", 971, 60], + ["Masterpiece of watercolor painting in Van Gogh style", "cat.png", 37890, 40] + ], [positive_input, i2i_input, seed_input, steps_input]) + + # if you are launching remotely, specify server_name and server_port + # demo.launch(server_name='your server name', server_port='server port in int') + # Read more in the docs: https://gradio.app/docs/ + # if you want create public link for sharing demo, please add share=True + demo.launch() + + +.. parsed-literal:: + + Running on local URL: http://127.0.0.1:7860 + + To create a public link, set `share=True` in `launch()`. + + + +.. raw:: html + +
+ + +.. code:: ipython3 + + demo.close() + del image2image_pipe + gc.collect() + + +.. parsed-literal:: + + Closing server running on port: 7860 + + + + +.. parsed-literal:: + + 312 + + + +SDXL Refiner model\ `⇑ <#top>`__ +############################################################################################################################### + + +As we discussed above, Stable Diffusion XL can be used in a 2-stages +approach: first, the base model is used to generate latents of the +desired output size. In the second step, we use a specialized +high-resolution model for the refinement of latents generated in the +first step, using the same prompt. The Stable Diffusion XL Refiner model +is designed to transform regular images into stunning masterpieces with +the help of user-specified prompt text. It can be used to improve the +quality of image generation after the Stable Diffusion XL Base. The +refiner model accepts latents produced by the SDXL base model and text +prompt for improving generated image. + +.. code:: ipython3 + + from optimum.intel import OVStableDiffusionXLImg2ImgPipeline, OVStableDiffusionXLPipeline + from pathlib import Path + + refiner_model_id = "stabilityai/stable-diffusion-xl-refiner-1.0" + refiner_model_dir = Path("openvino-sd-xl-refiner-1.0") + + + if not refiner_model_dir.exists(): + refiner = OVStableDiffusionXLImg2ImgPipeline.from_pretrained(refiner_model_id, export=True, compile=False) + refiner.half() + refiner.save_pretrained(refiner_model_dir) + del refiner + gc.collect() + +Select inference device\ `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=2, options=('CPU', 'GPU', 'AUTO'), value='AUTO') + + + +Run Text2Image generation with Refinement\ `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +.. code:: ipython3 + + import numpy as np + import gc + model_dir = Path("openvino-sd-xl-base-1.0") + base = OVStableDiffusionXLPipeline.from_pretrained(model_dir, device=device.value) + prompt = "cute cat 4k, high-res, masterpiece, best quality, soft lighting, dynamic angle" + latents = base(prompt, num_inference_steps=15, height=512, width=512, generator=np.random.RandomState(314), output_type="latent").images[0] + + del base + gc.collect() + + +.. parsed-literal:: + + Compiling the vae_decoder... + Compiling the unet... + Compiling the text_encoder_2... + Compiling the vae_encoder... + Compiling the text_encoder... + /home/ea/work/ov_notebooks_env/lib/python3.8/site-packages/optimum/intel/openvino/modeling_diffusion.py:552: FutureWarning: `shared_memory` is deprecated and will be removed in 2024.0. Value of `shared_memory` is going to override `share_inputs` value. Please use only `share_inputs` explicitly. + outputs = self.request(inputs, shared_memory=True) + + + +.. parsed-literal:: + + 0%| | 0/15 [00:00 +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/248-stable-diffusion-xl-with-output_files/ + +

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/248-stable-diffusion-xl-with-output_files/


../
+248-stable-diffusion-xl-with-output_10_3.jpg       16-Aug-2023 01:31               21518
+248-stable-diffusion-xl-with-output_10_3.png       16-Aug-2023 01:31              439606
+248-stable-diffusion-xl-with-output_18_3.jpg       16-Aug-2023 01:31               22630
+248-stable-diffusion-xl-with-output_18_3.png       16-Aug-2023 01:31              448218
+248-stable-diffusion-xl-with-output_29_3.jpg       16-Aug-2023 01:31               29603
+248-stable-diffusion-xl-with-output_29_3.png       16-Aug-2023 01:31              454022
+

+ diff --git a/docs/notebooks/249-oneformer-segmentation-with-output.rst b/docs/notebooks/249-oneformer-segmentation-with-output.rst new file mode 100644 index 00000000000..905f00c17ec --- /dev/null +++ b/docs/notebooks/249-oneformer-segmentation-with-output.rst @@ -0,0 +1,418 @@ +Universal Segmentation with OneFormer and OpenVINO +================================================== + +This tutorial demonstrates how to use the +`OneFormer `__ model from HuggingFace +with OpenVINO. It describes how to download weights and create PyTorch +model using Hugging Face transformers library, then convert model to +OpenVINO Intermediate Representation format (IR) using OpenVINO Model +Optimizer API and run model inference + +|image0| + +OneFormer is a follow-up work of +`Mask2Former `__. The latter still +requires training on instance/semantic/panoptic datasets separately to +get state-of-the-art results. + +OneFormer incorporates a text module in the Mask2Former framework, to +condition the model on the respective subtask (instance, semantic or +panoptic). This gives even more accurate results, but comes with a cost +of increased latency, however. + +.. |image0| image:: https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/oneformer_architecture.png + +Install required libraries +-------------------------- + +.. code:: ipython3 + + !pip install -q "transformers>=4.26.0" "openvino==2023.1.0.dev20230728" gradio torch scipy ipywidgets Pillow matplotlib + +Prepare the environment +----------------------- + +Import all required packages and set paths for models and constant +variables. + +.. code:: ipython3 + + import warnings + from collections import defaultdict + from pathlib import Path + import sys + + from transformers import OneFormerProcessor, OneFormerForUniversalSegmentation + from transformers.models.oneformer.modeling_oneformer import OneFormerForUniversalSegmentationOutput + import torch + import matplotlib.pyplot as plt + import matplotlib.patches as mpatches + from PIL import Image + from PIL import ImageOps + + import openvino + + sys.path.append("../utils") + from notebook_utils import download_file + + +.. parsed-literal:: + + 2023-08-13 20:13:13.033722: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-08-13 20:13:13.205781: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. + 2023-08-13 20:13:14.052205: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + + +.. code:: ipython3 + + IR_PATH = Path("oneformer.xml") + OUTPUT_NAMES = ['class_queries_logits', 'masks_queries_logits'] + +Load OneFormer fine-tuned on COCO for universal segmentation +------------------------------------------------------------ + +Here we use the ``from_pretrained`` method of +``OneFormerForUniversalSegmentation`` to load the `HuggingFace OneFormer +model `__ +based on Swin-L backbone and trained on +`COCO `__ dataset. + +Also, we use HuggingFace processor to prepare the model inputs from +images and post-process model outputs for visualization. + +.. code:: ipython3 + + processor = OneFormerProcessor.from_pretrained("shi-labs/oneformer_coco_swin_large") + model = OneFormerForUniversalSegmentation.from_pretrained( + "shi-labs/oneformer_coco_swin_large", + ) + id2label = model.config.id2label + +.. code:: ipython3 + + task_seq_length = processor.task_seq_length + shape = (800, 800) + dummy_input = { + "pixel_values": torch.randn(1, 3, *shape), + "task_inputs": torch.randn(1, task_seq_length), + "pixel_mask": torch.randn(1, *shape), + } + +Convert the model to OpenVINO IR format +--------------------------------------- + +Convert the PyTorch model to IR format to take advantage of OpenVINO +optimization tools and features. The ``openvino.convert_model`` python +function in OpenVINO Converter can convert the model. The function +returns instance of OpenVINO Model class, which is ready to use in +Python interface. However, it can also be serialized to OpenVINO IR +format for future execution using ``save_model`` function. PyTorch to +OpenVINO conversion is based on TorchScript tracing. HuggingFace models +have specific configuration parameter ``torchscript``, which can be used +for making the model more suitable for tracing. For preparing model. we +should provide PyTorch model instance and example input to +``openvino.convert_model``. + +.. code:: ipython3 + + model.config.torchscript = True + + if not IR_PATH.exists(): + with warnings.catch_warnings(): + warnings.simplefilter("ignore") + model = openvino.convert_model(model, example_input=dummy_input) + openvino.save_model(model, IR_PATH, compress_to_fp16=False) + +Select inference device +----------------------- + +Select device from dropdown list for running inference using OpenVINO + +.. code:: ipython3 + + import ipywidgets as widgets + + core = openvino.Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=4, options=('CPU', 'GPU.0', 'GPU.1', 'GPU.2', 'AUTO'), value='AUTO') + + + +We can prepare the image using the HuggingFace processor. OneFormer +leverages a processor which internally consists of an image processor +(for the image modality) and a tokenizer (for the text modality). +OneFormer is actually a multimodal model, since it incorporates both +images and text to solve image segmentation. + +.. code:: ipython3 + + def prepare_inputs(image: Image.Image, task: str): + """Convert image to model input""" + image = ImageOps.pad(image, shape) + inputs = processor(image, [task], return_tensors="pt") + converted = { + 'pixel_values': inputs['pixel_values'], + 'task_inputs': inputs['task_inputs'] + } + return converted + +.. code:: ipython3 + + def process_output(d): + """Convert OpenVINO model output to HuggingFace representation for visualization""" + hf_kwargs = { + output_name: torch.tensor(d[output_name]) for output_name in OUTPUT_NAMES + } + + return OneFormerForUniversalSegmentationOutput(**hf_kwargs) + +.. code:: ipython3 + + # Read the model from files. + model = core.read_model(model=IR_PATH) + # Compile the model. + compiled_model = core.compile_model(model=model, device_name=device.value) + +Model predicts ``class_queries_logits`` of shape +``(batch_size, num_queries)`` and ``masks_queries_logits`` of shape +``(batch_size, num_queries, height, width)``. + +Here we define functions for visualization of network outputs to show +the inference results. + +.. code:: ipython3 + + class Visualizer: + @staticmethod + def extract_legend(handles): + fig = plt.figure() + fig.legend(handles=handles, ncol=len(handles) // 20 + 1, loc='center') + fig.tight_layout() + return fig + + @staticmethod + def predicted_semantic_map_to_figure(predicted_map): + segmentation = predicted_map[0] + # get the used color map + viridis = plt.get_cmap('viridis', torch.max(segmentation)) + # get all the unique numbers + labels_ids = torch.unique(segmentation).tolist() + fig, ax = plt.subplots() + ax.imshow(segmentation) + ax.set_axis_off() + handles = [] + for label_id in labels_ids: + label = id2label[label_id] + color = viridis(label_id) + handles.append(mpatches.Patch(color=color, label=label)) + fig_legend = Visualizer.extract_legend(handles=handles) + fig.tight_layout() + return fig, fig_legend + + @staticmethod + def predicted_instance_map_to_figure(predicted_map): + segmentation = predicted_map[0]['segmentation'] + segments_info = predicted_map[0]['segments_info'] + # get the used color map + viridis = plt.get_cmap('viridis', torch.max(segmentation)) + fig, ax = plt.subplots() + ax.imshow(segmentation) + ax.set_axis_off() + instances_counter = defaultdict(int) + handles = [] + # for each segment, draw its legend + for segment in segments_info: + segment_id = segment['id'] + segment_label_id = segment['label_id'] + segment_label = id2label[segment_label_id] + label = f"{segment_label}-{instances_counter[segment_label_id]}" + instances_counter[segment_label_id] += 1 + color = viridis(segment_id) + handles.append(mpatches.Patch(color=color, label=label)) + + fig_legend = Visualizer.extract_legend(handles) + fig.tight_layout() + return fig, fig_legend + + @staticmethod + def predicted_panoptic_map_to_figure(predicted_map): + segmentation = predicted_map[0]['segmentation'] + segments_info = predicted_map[0]['segments_info'] + # get the used color map + viridis = plt.get_cmap('viridis', torch.max(segmentation)) + fig, ax = plt.subplots() + ax.imshow(segmentation) + ax.set_axis_off() + instances_counter = defaultdict(int) + handles = [] + # for each segment, draw its legend + for segment in segments_info: + segment_id = segment['id'] + segment_label_id = segment['label_id'] + segment_label = id2label[segment_label_id] + label = f"{segment_label}-{instances_counter[segment_label_id]}" + instances_counter[segment_label_id] += 1 + color = viridis(segment_id) + handles.append(mpatches.Patch(color=color, label=label)) + + fig_legend = Visualizer.extract_legend(handles) + fig.tight_layout() + return fig, fig_legend + +.. code:: ipython3 + + def segment(img: Image.Image, task: str): + """ + Apply segmentation on an image. + + Args: + img: Input image. It will be resized to 800x800. + task: String describing the segmentation task. Supported values are: "semantic", "instance" and "panoptic". + Returns: + Tuple[Figure, Figure]: Segmentation map and legend charts. + """ + if img is None: + raise gr.Error("Please load the image or use one from the examples list") + inputs = prepare_inputs(img, task) + outputs = compiled_model(inputs) + hf_output = process_output(outputs) + predicted_map = getattr(processor, f"post_process_{task}_segmentation")( + hf_output, target_sizes=[img.size[::-1]] + ) + return getattr(Visualizer, f"predicted_{task}_map_to_figure")(predicted_map) + +.. code:: ipython3 + + image = download_file("http://images.cocodataset.org/val2017/000000439180.jpg", "sample.jpg") + image = Image.open("sample.jpg") + image + + + +.. parsed-literal:: + + sample.jpg: 0%| | 0.00/194k [00:00 +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/249-oneformer-segmentation-with-output_files/ + +

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/249-oneformer-segmentation-with-output_files/


../
+249-oneformer-segmentation-with-output_22_1.jpg    16-Aug-2023 01:31               64470
+249-oneformer-segmentation-with-output_22_1.png    16-Aug-2023 01:31              514894
+249-oneformer-segmentation-with-output_26_0.jpg    16-Aug-2023 01:31               29211
+249-oneformer-segmentation-with-output_26_0.png    16-Aug-2023 01:31              172371
+

+ diff --git a/docs/notebooks/301-tensorflow-training-openvino-nncf-with-output.rst b/docs/notebooks/301-tensorflow-training-openvino-nncf-with-output.rst index 51bfe312307..cb6a4eb137d 100644 --- a/docs/notebooks/301-tensorflow-training-openvino-nncf-with-output.rst +++ b/docs/notebooks/301-tensorflow-training-openvino-nncf-with-output.rst @@ -6,11 +6,26 @@ created in `301-tensorflow-training-openvino.ipynb <301-tensorflow-training-openvino.ipynb>`__, to improve inference speed. Quantization is performed with `Post-training Quantization with -NNCF `__. +NNCF `__. A custom dataloader and metric will be defined, and accuracy and performance will be computed for the original IR model and the quantized model. +**Table of contents**: + +- `Preparation <#preparation>`__ + + - `Imports <#imports>`__ + +- `Post-training Quantization with NNCF <#post-training-quantization-with-nncf>`__ + + - `Select inference device <#post-training-quantization-with-nncf>`__ + +- `Compare Metrics <#post-training-quantization-with-nncf>`__ +- `Run Inference on Quantized Model <#run-inference-on-quantized-model>`__ +- `Compare Inference Speed <#compare-inference-speed>`__ + + Preparation ----------- @@ -38,10 +53,10 @@ notebook. This will take a while. .. parsed-literal:: - 2023-07-11 23:45:46.901183: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. - 2023-07-11 23:45:46.936281: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + 2023-07-05 23:54:28.962752: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-07-05 23:54:28.997784: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. - 2023-07-11 23:45:47.523785: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + 2023-07-05 23:54:29.609276: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT .. parsed-literal:: @@ -54,7 +69,7 @@ notebook. This will take a while. .. parsed-literal:: - 2023-07-11 23:45:49.130715: W tensorflow/core/common_runtime/gpu/gpu_device.cc:1956] Cannot dlopen some GPU libraries. Please make sure the missing libraries mentioned above are installed properly if you would like to use GPU. Follow the guide at https://www.tensorflow.org/install/gpu for how to download and setup the required libraries for your platform. + 2023-07-05 23:54:31.178171: W tensorflow/core/common_runtime/gpu/gpu_device.cc:1956] Cannot dlopen some GPU libraries. Please make sure the missing libraries mentioned above are installed properly if you would like to use GPU. Follow the guide at https://www.tensorflow.org/install/gpu for how to download and setup the required libraries for your platform. Skipping registering GPU devices... @@ -67,9 +82,9 @@ notebook. This will take a while. .. parsed-literal:: - 2023-07-11 23:45:49.466286: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] + 2023-07-05 23:54:31.493885: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] [[{{node Placeholder/_4}}]] - 2023-07-11 23:45:49.466548: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] + 2023-07-05 23:54:31.494167: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] [[{{node Placeholder/_4}}]] @@ -79,28 +94,28 @@ notebook. This will take a while. .. parsed-literal:: - 2023-07-11 23:45:49.938600: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] + 2023-07-05 23:54:31.947372: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] [[{{node Placeholder/_4}}]] - 2023-07-11 23:45:49.938843: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] + 2023-07-05 23:54:31.947613: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] [[{{node Placeholder/_4}}]] - 2023-07-11 23:45:50.062635: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] + 2023-07-05 23:54:32.077841: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] + [[{{node Placeholder/_4}}]] + 2023-07-05 23:54:32.078164: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] [[{{node Placeholder/_4}}]] - 2023-07-11 23:45:50.062905: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [2936] - [[{{node Placeholder/_0}}]] .. parsed-literal:: (32, 180, 180, 3) (32,) - 0.0 0.9967369 + 0.0 1.0 .. parsed-literal:: - 2023-07-11 23:45:50.838129: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] - [[{{node Placeholder/_4}}]] - 2023-07-11 23:45:50.838521: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] + 2023-07-05 23:54:32.897047: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [2936] + [[{{node Placeholder/_0}}]] + 2023-07-05 23:54:32.897375: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] [[{{node Placeholder/_4}}]] @@ -151,55 +166,55 @@ notebook. This will take a while. .. parsed-literal:: - 2023-07-11 23:45:51.781311: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] + 2023-07-05 23:54:33.773069: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [2936] + [[{{node Placeholder/_0}}]] + 2023-07-05 23:54:33.773519: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] [[{{node Placeholder/_4}}]] - 2023-07-11 23:45:51.781738: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [2936] + + +.. parsed-literal:: + + 92/92 [==============================] - ETA: 0s - loss: 1.2943 - accuracy: 0.4486 + +.. parsed-literal:: + + 2023-07-05 23:54:40.025734: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [734] + [[{{node Placeholder/_0}}]] + 2023-07-05 23:54:40.026032: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [734] [[{{node Placeholder/_0}}]] .. parsed-literal:: - 92/92 [==============================] - ETA: 0s - loss: 1.3257 - accuracy: 0.4353 - -.. parsed-literal:: - - 2023-07-11 23:45:58.082635: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [734] - [[{{node Placeholder/_4}}]] - 2023-07-11 23:45:58.082922: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [734] - [[{{node Placeholder/_4}}]] - - -.. parsed-literal:: - - 92/92 [==============================] - 7s 66ms/step - loss: 1.3257 - accuracy: 0.4353 - val_loss: 1.1364 - val_accuracy: 0.5341 + 92/92 [==============================] - 7s 66ms/step - loss: 1.2943 - accuracy: 0.4486 - val_loss: 1.0944 - val_accuracy: 0.5354 Epoch 2/15 - 92/92 [==============================] - 6s 63ms/step - loss: 1.0419 - accuracy: 0.5872 - val_loss: 1.0635 - val_accuracy: 0.5886 + 92/92 [==============================] - 6s 63ms/step - loss: 1.0396 - accuracy: 0.5787 - val_loss: 0.9602 - val_accuracy: 0.6322 Epoch 3/15 - 92/92 [==============================] - 6s 63ms/step - loss: 0.9311 - accuracy: 0.6352 - val_loss: 0.9998 - val_accuracy: 0.6131 + 92/92 [==============================] - 6s 64ms/step - loss: 0.9646 - accuracy: 0.6213 - val_loss: 0.9223 - val_accuracy: 0.6417 Epoch 4/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.8423 - accuracy: 0.6788 - val_loss: 0.9286 - val_accuracy: 0.6703 + 92/92 [==============================] - 6s 64ms/step - loss: 0.8775 - accuracy: 0.6533 - val_loss: 0.8511 - val_accuracy: 0.6594 Epoch 5/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.8117 - accuracy: 0.6832 - val_loss: 0.8297 - val_accuracy: 0.6567 + 92/92 [==============================] - 6s 64ms/step - loss: 0.8354 - accuracy: 0.6884 - val_loss: 0.8471 - val_accuracy: 0.6689 Epoch 6/15 - 92/92 [==============================] - 6s 63ms/step - loss: 0.7639 - accuracy: 0.7078 - val_loss: 0.7671 - val_accuracy: 0.7112 + 92/92 [==============================] - 6s 64ms/step - loss: 0.7722 - accuracy: 0.7033 - val_loss: 0.8405 - val_accuracy: 0.6935 Epoch 7/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.7212 - accuracy: 0.7204 - val_loss: 0.7973 - val_accuracy: 0.6962 + 92/92 [==============================] - 6s 64ms/step - loss: 0.7347 - accuracy: 0.7207 - val_loss: 0.8848 - val_accuracy: 0.6730 Epoch 8/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.6783 - accuracy: 0.7435 - val_loss: 0.8085 - val_accuracy: 0.6717 + 92/92 [==============================] - 6s 63ms/step - loss: 0.6980 - accuracy: 0.7469 - val_loss: 0.7724 - val_accuracy: 0.6948 Epoch 9/15 - 92/92 [==============================] - 6s 63ms/step - loss: 0.6554 - accuracy: 0.7500 - val_loss: 0.7403 - val_accuracy: 0.7193 + 92/92 [==============================] - 6s 64ms/step - loss: 0.6629 - accuracy: 0.7476 - val_loss: 0.7512 - val_accuracy: 0.7071 Epoch 10/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.6283 - accuracy: 0.7629 - val_loss: 0.6977 - val_accuracy: 0.7153 + 92/92 [==============================] - 6s 63ms/step - loss: 0.6429 - accuracy: 0.7643 - val_loss: 0.7196 - val_accuracy: 0.7125 Epoch 11/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.6220 - accuracy: 0.7592 - val_loss: 0.7095 - val_accuracy: 0.7343 + 92/92 [==============================] - 6s 64ms/step - loss: 0.5967 - accuracy: 0.7755 - val_loss: 0.7228 - val_accuracy: 0.7084 Epoch 12/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.5818 - accuracy: 0.7813 - val_loss: 0.7068 - val_accuracy: 0.7234 + 92/92 [==============================] - 6s 63ms/step - loss: 0.5860 - accuracy: 0.7769 - val_loss: 0.7501 - val_accuracy: 0.7153 Epoch 13/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.5503 - accuracy: 0.7922 - val_loss: 0.6985 - val_accuracy: 0.7207 + 92/92 [==============================] - 6s 64ms/step - loss: 0.5695 - accuracy: 0.7793 - val_loss: 0.7366 - val_accuracy: 0.7153 Epoch 14/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.5497 - accuracy: 0.7953 - val_loss: 0.6983 - val_accuracy: 0.7439 + 92/92 [==============================] - 6s 63ms/step - loss: 0.5392 - accuracy: 0.7970 - val_loss: 0.7375 - val_accuracy: 0.7275 Epoch 15/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.5285 - accuracy: 0.8028 - val_loss: 0.7165 - val_accuracy: 0.7425 + 92/92 [==============================] - 6s 64ms/step - loss: 0.5098 - accuracy: 0.8048 - val_loss: 0.6984 - val_accuracy: 0.7330 @@ -209,46 +224,46 @@ notebook. This will take a while. .. parsed-literal:: 1/1 [==============================] - 0s 76ms/step - This image most likely belongs to sunflowers with a 92.57 percent confidence. + This image most likely belongs to sunflowers with a 99.23 percent confidence. .. parsed-literal:: - 2023-07-11 23:47:21.447950: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'random_flip_input' with dtype float and shape [?,180,180,3] + 2023-07-05 23:56:03.289411: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'random_flip_input' with dtype float and shape [?,180,180,3] [[{{node random_flip_input}}]] - 2023-07-11 23:47:21.533588: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-07-05 23:56:03.376040: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:47:21.543460: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'random_flip_input' with dtype float and shape [?,180,180,3] + 2023-07-05 23:56:03.385907: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'random_flip_input' with dtype float and shape [?,180,180,3] [[{{node random_flip_input}}]] - 2023-07-11 23:47:21.555445: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-07-05 23:56:03.396762: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:47:21.562382: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-07-05 23:56:03.403700: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:47:21.569165: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-07-05 23:56:03.410703: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:47:21.579854: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-07-05 23:56:03.421394: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:47:21.618942: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'sequential_1_input' with dtype float and shape [?,180,180,3] + 2023-07-05 23:56:03.461681: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'sequential_1_input' with dtype float and shape [?,180,180,3] [[{{node sequential_1_input}}]] - 2023-07-11 23:47:21.686887: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-07-05 23:56:03.529355: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:47:21.707365: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'sequential_1_input' with dtype float and shape [?,180,180,3] + 2023-07-05 23:56:03.549619: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'sequential_1_input' with dtype float and shape [?,180,180,3] [[{{node sequential_1_input}}]] - 2023-07-11 23:47:21.746351: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,22,22,64] + 2023-07-05 23:56:03.588567: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,22,22,64] [[{{node inputs}}]] - 2023-07-11 23:47:21.769890: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-07-05 23:56:03.611996: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:47:21.843307: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-07-05 23:56:03.685894: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:47:21.985172: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-07-05 23:56:03.828047: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:47:22.122564: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,22,22,64] + 2023-07-05 23:56:03.965814: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,22,22,64] [[{{node inputs}}]] - 2023-07-11 23:47:22.156382: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-07-05 23:56:03.999799: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:47:22.184461: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-07-05 23:56:04.028229: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:47:22.230863: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-07-05 23:56:04.074705: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] WARNING:absl:Found untraced functions such as _jit_compiled_convolution_op, _jit_compiled_convolution_op, _jit_compiled_convolution_op, _update_step_xla while saving (showing 4 of 4). These functions will not be directly callable after loading. @@ -273,7 +288,7 @@ notebook. This will take a while. (1, 180, 180, 3) [1,180,180,3] - This image most likely belongs to dandelion with a 95.29 percent confidence. + This image most likely belongs to dandelion with a 99.81 percent confidence. @@ -316,10 +331,8 @@ OpenVINO with minimal accuracy drop. Create a quantized model from the pre-trained FP32 model and the calibration dataset. The optimization process contains the following -steps: - -1. Create a Dataset for quantization. -2. Run nncf.quantize for getting an optimized model. +steps: 1. Create a Dataset for quantization. 2. Run nncf.quantize for +getting an optimized model. The validation dataset already defined in the training notebook. @@ -350,9 +363,9 @@ The validation dataset already defined in the training notebook. .. parsed-literal:: - 2023-07-11 23:47:25.157142: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [734] + 2023-07-05 23:56:07.075279: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [734] [[{{node Placeholder/_4}}]] - 2023-07-11 23:47:25.157516: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [734] + 2023-07-05 23:56:07.075533: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [734] [[{{node Placeholder/_4}}]] @@ -399,8 +412,8 @@ control ` .. parsed-literal:: - Statistics collection: 73%|███████▎ | 734/1000 [00:04<00:01, 163.35it/s] - Biases correction: 100%|██████████| 5/5 [00:01<00:00, 3.98it/s] + Statistics collection: 73%|███████▎ | 734/1000 [00:04<00:01, 166.65it/s] + Biases correction: 100%|██████████| 5/5 [00:01<00:00, 3.99it/s] Save quantized model to benchmark. @@ -463,8 +476,8 @@ Calculate accuracy for the original model and the quantized model. .. parsed-literal:: - Accuracy of the original model: 0.743 - Accuracy of the quantized model: 0.741 + Accuracy of the original model: 0.733 + Accuracy of the quantized model: 0.737 Compare file size of the models. @@ -556,7 +569,7 @@ Python API. 'output/A_Close_Up_Photo_of_a_Dandelion.jpg' already exists. input image shape: (1, 180, 180, 3) input layer shape: [1,180,180,3] - This image most likely belongs to dandelion with a 95.55 percent confidence. + This image most likely belongs to dandelion with a 99.82 percent confidence. @@ -630,7 +643,7 @@ measured for CPU+GPU as well. The number of seconds is set to 15. [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.THROUGHPUT. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 11.63 ms + [ INFO ] Read model took 12.02 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] sequential_1_input (node: sequential_1_input) : f32 / [...] / [1,180,180,3] @@ -644,7 +657,7 @@ measured for CPU+GPU as well. The number of seconds is set to 15. [ INFO ] Model outputs: [ INFO ] outputs (node: sequential_2/outputs/BiasAdd) : f32 / [...] / [1,5] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 68.32 ms + [ INFO ] Compile model took 76.79 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: [ INFO ] NETWORK_NAME: TensorFlow_Frontend_IR @@ -666,17 +679,17 @@ measured for CPU+GPU as well. The number of seconds is set to 15. [ INFO ] Fill input 'sequential_1_input' with random values [Step 10/11] Measuring performance (Start inference asynchronously, 12 inference requests, limits: 15000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 7.07 ms + [ INFO ] First inference took 7.22 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 57672 iterations - [ INFO ] Duration: 15004.09 ms + [ INFO ] Count: 57276 iterations + [ INFO ] Duration: 15002.57 ms [ INFO ] Latency: - [ INFO ] Median: 2.92 ms - [ INFO ] Average: 2.94 ms - [ INFO ] Min: 1.83 ms - [ INFO ] Max: 10.49 ms - [ INFO ] Throughput: 3843.75 FPS + [ INFO ] Median: 2.90 ms + [ INFO ] Average: 2.95 ms + [ INFO ] Min: 1.67 ms + [ INFO ] Max: 234.29 ms + [ INFO ] Throughput: 3817.75 FPS .. code:: ipython3 @@ -702,7 +715,7 @@ measured for CPU+GPU as well. The number of seconds is set to 15. [ WARNING ] Performance hint was not explicitly specified in command line. Device(CPU) performance hint will be set to PerformanceMode.THROUGHPUT. [Step 4/11] Reading model files [ INFO ] Loading model files - [ INFO ] Read model took 13.24 ms + [ INFO ] Read model took 12.35 ms [ INFO ] Original model I/O parameters: [ INFO ] Model inputs: [ INFO ] sequential_1_input (node: sequential_1_input) : f32 / [...] / [1,180,180,3] @@ -716,7 +729,7 @@ measured for CPU+GPU as well. The number of seconds is set to 15. [ INFO ] Model outputs: [ INFO ] outputs (node: sequential_2/outputs/BiasAdd) : f32 / [...] / [1,5] [Step 7/11] Loading the model to the device - [ INFO ] Compile model took 56.10 ms + [ INFO ] Compile model took 54.95 ms [Step 8/11] Querying optimal runtime parameters [ INFO ] Model: [ INFO ] NETWORK_NAME: TensorFlow_Frontend_IR @@ -738,17 +751,17 @@ measured for CPU+GPU as well. The number of seconds is set to 15. [ INFO ] Fill input 'sequential_1_input' with random values [Step 10/11] Measuring performance (Start inference asynchronously, 12 inference requests, limits: 15000 ms duration) [ INFO ] Benchmarking in inference only mode (inputs filling are not included in measurement loop). - [ INFO ] First inference took 1.97 ms + [ INFO ] First inference took 2.06 ms [Step 11/11] Dumping statistics report [ INFO ] Execution Devices:['CPU'] - [ INFO ] Count: 179064 iterations - [ INFO ] Duration: 15001.51 ms + [ INFO ] Count: 178752 iterations + [ INFO ] Duration: 15001.22 ms [ INFO ] Latency: [ INFO ] Median: 0.92 ms [ INFO ] Average: 0.92 ms - [ INFO ] Min: 0.55 ms - [ INFO ] Max: 4.95 ms - [ INFO ] Throughput: 11936.40 FPS + [ INFO ] Min: 0.54 ms + [ INFO ] Max: 4.90 ms + [ INFO ] Throughput: 11915.83 FPS **Benchmark on MULTI:CPU,GPU** @@ -818,14 +831,14 @@ cached to the ``model_cache`` directory. .. parsed-literal:: - [ INFO ] Count: 58656 iterations - [ INFO ] Duration: 15003.30 ms + [ INFO ] Count: 58332 iterations + [ INFO ] Duration: 15005.08 ms [ INFO ] Latency: [ INFO ] Median: 2.88 ms - [ INFO ] Average: 2.88 ms - [ INFO ] Min: 1.26 ms - [ INFO ] Max: 11.15 ms - [ INFO ] Throughput: 3909.54 FPS + [ INFO ] Average: 2.89 ms + [ INFO ] Min: 2.02 ms + [ INFO ] Max: 8.94 ms + [ INFO ] Throughput: 3887.48 FPS **Quantized IR model - CPU** @@ -840,14 +853,14 @@ cached to the ``model_cache`` directory. .. parsed-literal:: - [ INFO ] Count: 179484 iterations - [ INFO ] Duration: 15001.08 ms + [ INFO ] Count: 179124 iterations + [ INFO ] Duration: 15001.17 ms [ INFO ] Latency: [ INFO ] Median: 0.92 ms [ INFO ] Average: 0.92 ms [ INFO ] Min: 0.56 ms - [ INFO ] Max: 4.81 ms - [ INFO ] Throughput: 11964.74 FPS + [ INFO ] Max: 4.33 ms + [ INFO ] Throughput: 11940.67 FPS **Original IR model - MULTI:CPU,GPU** diff --git a/docs/notebooks/301-tensorflow-training-openvino-nncf-with-output_files/301-tensorflow-training-openvino-nncf-with-output_2_15.png b/docs/notebooks/301-tensorflow-training-openvino-nncf-with-output_files/301-tensorflow-training-openvino-nncf-with-output_2_15.png index e912768cd43..236759738b3 100644 --- a/docs/notebooks/301-tensorflow-training-openvino-nncf-with-output_files/301-tensorflow-training-openvino-nncf-with-output_2_15.png +++ b/docs/notebooks/301-tensorflow-training-openvino-nncf-with-output_files/301-tensorflow-training-openvino-nncf-with-output_2_15.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:3ade1ae22d657137ea11d2ac028b436a5bb8bb354fe4585dea7102f45e14c013 -size 56504 +oid sha256:e3af8d92a72b6dfb54c116ab9a052ffc4d6783058db964f3b9ac4c63eb246a4f +size 55498 diff --git a/docs/notebooks/301-tensorflow-training-openvino-nncf-with-output_files/301-tensorflow-training-openvino-nncf-with-output_2_9.png b/docs/notebooks/301-tensorflow-training-openvino-nncf-with-output_files/301-tensorflow-training-openvino-nncf-with-output_2_9.png index 7739333a34e..eeb03a11926 100644 --- a/docs/notebooks/301-tensorflow-training-openvino-nncf-with-output_files/301-tensorflow-training-openvino-nncf-with-output_2_9.png +++ b/docs/notebooks/301-tensorflow-training-openvino-nncf-with-output_files/301-tensorflow-training-openvino-nncf-with-output_2_9.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:54ed774a58a533ff0c28ccdd8e02697dd11b799f992e201bdfa6082553bc8eb1 -size 585760 +oid sha256:a62d60162298ff48fa9224f12ef371e8a62ec7778c10eadca436339f74aa9253 +size 433486 diff --git a/docs/notebooks/301-tensorflow-training-openvino-nncf-with-output_files/index.html b/docs/notebooks/301-tensorflow-training-openvino-nncf-with-output_files/index.html deleted file mode 100644 index 0ac08c8d3f5..00000000000 --- a/docs/notebooks/301-tensorflow-training-openvino-nncf-with-output_files/index.html +++ /dev/null @@ -1,11 +0,0 @@ - -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/301-tensorflow-training-openvino-nncf-with-output_files/ - -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/301-tensorflow-training-openvino-nncf-with-output_files/


../
-301-tensorflow-training-openvino-nncf-with-outp..> 12-Jul-2023 00:11              143412
-301-tensorflow-training-openvino-nncf-with-outp..> 12-Jul-2023 00:11               56504
-301-tensorflow-training-openvino-nncf-with-outp..> 12-Jul-2023 00:11              143412
-301-tensorflow-training-openvino-nncf-with-outp..> 12-Jul-2023 00:11              941151
-301-tensorflow-training-openvino-nncf-with-outp..> 12-Jul-2023 00:11              585760
-

- diff --git a/docs/notebooks/301-tensorflow-training-openvino-with-output.rst b/docs/notebooks/301-tensorflow-training-openvino-with-output.rst index ad4e1cb2e06..a4f82b05b1f 100644 --- a/docs/notebooks/301-tensorflow-training-openvino-with-output.rst +++ b/docs/notebooks/301-tensorflow-training-openvino-with-output.rst @@ -1,6 +1,39 @@ From Training to Deployment with TensorFlow and OpenVINO™ ========================================================= +.. _top: + +**Table of contents**: + +- `TensorFlow Image Classification Training <#tensorflow-image-classification-training>`__ +- `Import TensorFlow and Other Libraries <#import-tensorflow-and-other-libraries>`__ +- `Download and Explore the Dataset <#download-and-explore-the-dataset>`__ +- `Load Using keras.preprocessing <#load-using-keras.preprocessing>`__ +- `Create a Dataset <#create-a-dataset>`__ +- `Visualize the Data <#visualize-the-data>`__ +- `Configure the Dataset for Performance <#configure-the-dataset-for-performance>`__ +- `Standardize the Data <#standardize-the-data>`__ +- `Create the Model <#create-the-model>`__ +- `Compile the Model <#compile-the-model>`__ +- `Model Summary <#model-summary>`__ +- `Train the Model <#train-the-model>`__ +- `Visualize Training Results <#visualize-training-results>`__ +- `Overfitting <#overfitting>`__ +- `Data Augmentation <#data-augmentation>`__ +- `Dropout <#dropout>`__ +- `Compile and Train the Model <#compile-and-train-the-model>`__ +- `Visualize Training Results <#visualize-training-results>`__ +- `Predict on New Data <#predict-on-new-data>`__ +- `Save the TensorFlow Model <#save-the-tensorflow-model>`__ +- `Convert the TensorFlow model with OpenVINO Model Optimizer <#convert-the-tensorflow-model-with-openvino-model-optimizer>`__ +- `Preprocessing Image Function <#preprocessing-image-function>`__ +- `OpenVINO Runtime Setup <#openvino-runtime-setup>`__ + + - `Select inference device <#select-inference-device>`__ + +- `Run the Inference Step <#run-the-inference-step>`__ +- `The Next Steps <#the-next-steps>`__ + .. code:: ipython3 # @title Licensed under the Apache License, Version 2.0 (the "License"); @@ -22,8 +55,8 @@ From Training to Deployment with TensorFlow and OpenVINO™ This tutorial demonstrates how to train, convert, and deploy an image classification model with TensorFlow and OpenVINO. This particular notebook shows the process where we perform the inference step on the -freshly trained model that is converted to OpenVINO IR with Model -Optimizer. For faster inference speed on the model created in this +freshly trained model that is converted to OpenVINO IR with model +conversion API. For faster inference speed on the model created in this notebook, check out the `Post-Training Quantization with TensorFlow Classification Model <./301-tensorflow-training-openvino-nncf.ipynb>`__ notebook. @@ -33,12 +66,13 @@ Classification Tutorial `__ in its entirety. -The **flower_ir.bin** and **flower_ir.xml** (pre-trained models) can be -obtained by executing the code with ‘Runtime->Run All’ or the Ctrl+F9 -command. +The ``flower_ir.bin`` and ``flower_ir.xml`` (pre-trained models) can be +obtained by executing the code with ‘Runtime->Run All’ or the +``Ctrl+F9`` command. + +TensorFlow Image Classification Training `⇑ <#top>`__ +############################################################################################################################### -TensorFlow Image Classification Training ----------------------------------------- The first part of the tutorial shows how to classify images of flowers (based on the TensorFlow’s official tutorial). It creates an image @@ -58,8 +92,9 @@ This tutorial follows a basic machine learning workflow: 4. Train the model 5. Test the model -Import TensorFlow and Other Libraries -------------------------------------- +Import TensorFlow and Other Libraries `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -85,14 +120,15 @@ Import TensorFlow and Other Libraries .. parsed-literal:: - 2023-07-11 23:48:43.746139: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. - 2023-07-11 23:48:43.781112: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + 2023-08-16 01:08:54.169184: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-08-16 01:08:54.203604: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. - 2023-07-11 23:48:44.293875: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + 2023-08-16 01:08:54.707315: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT -Download and Explore the Dataset --------------------------------- +Download and Explore the Dataset `⇑ <#top>`__ +############################################################################################################################### + This tutorial uses a dataset of about 3,700 photos of flowers. The dataset contains 5 sub-directories, one per class: @@ -177,8 +213,9 @@ And some tulips: -Load Using keras.preprocessing ------------------------------- +Load Using keras.preprocessing `⇑ <#top>`__ +############################################################################################################################### + Let’s load these images off disk using the helpful `image_dataset_from_directory `__ @@ -188,8 +225,9 @@ also write your own data loading code from scratch by visiting the `load images `__ tutorial. -Create a Dataset ----------------- +Create a Dataset `⇑ <#top>`__ +############################################################################################################################### + Define some parameters for the loader: @@ -221,7 +259,7 @@ Let’s use 80% of the images for training, and 20% for validation. .. parsed-literal:: - 2023-07-11 23:48:45.680125: W tensorflow/core/common_runtime/gpu/gpu_device.cc:1956] Cannot dlopen some GPU libraries. Please make sure the missing libraries mentioned above are installed properly if you would like to use GPU. Follow the guide at https://www.tensorflow.org/install/gpu for how to download and setup the required libraries for your platform. + 2023-08-16 01:08:56.066599: W tensorflow/core/common_runtime/gpu/gpu_device.cc:1956] Cannot dlopen some GPU libraries. Please make sure the missing libraries mentioned above are installed properly if you would like to use GPU. Follow the guide at https://www.tensorflow.org/install/gpu for how to download and setup the required libraries for your platform. Skipping registering GPU devices... @@ -256,8 +294,9 @@ datasets. These correspond to the directory names in alphabetical order. ['daisy', 'dandelion', 'roses', 'sunflowers', 'tulips'] -Visualize the Data ------------------- +Visualize the Data `⇑ <#top>`__ +############################################################################################################################### + Here are the first 9 images from the training dataset. @@ -274,10 +313,10 @@ Here are the first 9 images from the training dataset. .. parsed-literal:: - 2023-07-11 23:48:46.046287: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] + 2023-08-16 01:08:56.428488: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] + [[{{node Placeholder/_4}}]] + 2023-08-16 01:08:56.429092: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] [[{{node Placeholder/_4}}]] - 2023-07-11 23:48:46.046887: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [2936] - [[{{node Placeholder/_0}}]] @@ -304,9 +343,9 @@ over the dataset and retrieve batches of images: .. parsed-literal:: - 2023-07-11 23:48:46.535781: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [2936] - [[{{node Placeholder/_0}}]] - 2023-07-11 23:48:46.536007: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [2936] + 2023-08-16 01:08:56.917347: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] + [[{{node Placeholder/_4}}]] + 2023-08-16 01:08:56.917776: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [2936] [[{{node Placeholder/_0}}]] @@ -318,8 +357,9 @@ shape ``(32,)``, these are corresponding labels to the 32 images. You can call ``.numpy()`` on the ``image_batch`` and ``labels_batch`` tensors to convert them to a ``numpy.ndarray``. -Configure the Dataset for Performance -------------------------------------- +Configure the Dataset for Performance `⇑ <#top>`__ +############################################################################################################################### + Let’s make sure to use buffered prefetching so you can yield data from disk without having I/O become blocking. These are two important methods @@ -344,8 +384,9 @@ guide `__. train_ds = train_ds.cache().shuffle(1000).prefetch(buffer_size=AUTOTUNE) val_ds = val_ds.cache().prefetch(buffer_size=AUTOTUNE) -Standardize the Data --------------------- +Standardize the Data `⇑ <#top>`__ +############################################################################################################################### + The RGB channel values are in the ``[0, 255]`` range. This is not ideal for a neural network; in general you should seek to make your input @@ -373,15 +414,15 @@ calling map: .. parsed-literal:: - 2023-07-11 23:48:46.736316: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] - [[{{node Placeholder/_4}}]] - 2023-07-11 23:48:46.736650: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [2936] + 2023-08-16 01:08:57.116807: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [2936] + [[{{node Placeholder/_0}}]] + 2023-08-16 01:08:57.117197: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [2936] [[{{node Placeholder/_0}}]] .. parsed-literal:: - 0.0 1.0 + 0.0 0.9891067 Or, you can include the layer inside your model definition, which can @@ -393,8 +434,9 @@ logic in your model as well, you can use the `Resizing `__ layer. -Create the Model ----------------- +Create the Model `⇑ <#top>`__ +############################################################################################################################### + The model consists of three convolution blocks with a max pool layer in each of them. There’s a fully connected layer with 128 units on top of @@ -419,8 +461,9 @@ standard approach. layers.Dense(num_classes) ]) -Compile the Model ------------------ +Compile the Model `⇑ <#top>`__ +############################################################################################################################### + For this tutorial, choose the ``optimizers.Adam`` optimizer and ``losses.SparseCategoricalCrossentropy`` loss function. To view training @@ -433,8 +476,9 @@ argument. loss=tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True), metrics=['accuracy']) -Model Summary -------------- +Model Summary `⇑ <#top>`__ +############################################################################################################################### + View all the layers of the network using the model’s ``summary`` method. @@ -445,8 +489,9 @@ View all the layers of the network using the model’s ``summary`` method. # model.summary() -Train the Model ---------------- +Train the Model `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -457,8 +502,9 @@ Train the Model # epochs=epochs # ) -Visualize Training Results --------------------------- +Visualize Training Results `⇑ <#top>`__ +############################################################################################################################### + Create plots of loss and accuracy on the training and validation sets. @@ -493,8 +539,9 @@ accuracy on the validation set. Let’s look at what went wrong and try to increase the overall performance of the model. -Overfitting ------------ +Overfitting `⇑ <#top>`__ +############################################################################################################################### + In the plots above, the training accuracy is increasing linearly over time, whereas validation accuracy stalls around 60% in the training @@ -512,8 +559,9 @@ There are multiple ways to fight overfitting in the training process. In this tutorial, you’ll use *data augmentation* and add *Dropout* to your model. -Data Augmentation ------------------ +Data Augmentation `⇑ <#top>`__ +############################################################################################################################### + Overfitting generally occurs when there are a small number of training examples. `Data @@ -556,10 +604,10 @@ augmentation to the same image several times: .. parsed-literal:: - 2023-07-11 23:48:47.665043: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] + 2023-08-16 01:08:57.956457: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] + [[{{node Placeholder/_4}}]] + 2023-08-16 01:08:57.956841: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] [[{{node Placeholder/_4}}]] - 2023-07-11 23:48:47.666032: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [2936] - [[{{node Placeholder/_0}}]] @@ -568,8 +616,9 @@ augmentation to the same image several times: You will use data augmentation to train a model in a moment. -Dropout -------- +Dropout `⇑ <#top>`__ +############################################################################################################################### + Another technique to reduce overfitting is to introduce `Dropout `__ @@ -601,8 +650,9 @@ it using augmented images. layers.Dense(num_classes, name="outputs") ]) -Compile and Train the Model ---------------------------- +Compile and Train the Model `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -672,59 +722,60 @@ Compile and Train the Model .. parsed-literal:: - 2023-07-11 23:48:48.689479: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] - [[{{node Placeholder/_4}}]] - 2023-07-11 23:48:48.689842: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] + 2023-08-16 01:08:58.847518: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [2936] + [[{{node Placeholder/_0}}]] + 2023-08-16 01:08:58.847798: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [2936] [[{{node Placeholder/_4}}]] .. parsed-literal:: - 92/92 [==============================] - ETA: 0s - loss: 1.2325 - accuracy: 0.4789 + 92/92 [==============================] - ETA: 0s - loss: 1.3880 - accuracy: 0.4196 .. parsed-literal:: - 2023-07-11 23:48:55.021722: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [734] - [[{{node Placeholder/_0}}]] - 2023-07-11 23:48:55.022051: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [734] + 2023-08-16 01:09:05.080237: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [734] [[{{node Placeholder/_0}}]] + 2023-08-16 01:09:05.080525: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int32 and shape [734] + [[{{node Placeholder/_4}}]] .. parsed-literal:: - 92/92 [==============================] - 7s 67ms/step - loss: 1.2325 - accuracy: 0.4789 - val_loss: 1.1056 - val_accuracy: 0.5681 + 92/92 [==============================] - 7s 65ms/step - loss: 1.3880 - accuracy: 0.4196 - val_loss: 1.1062 - val_accuracy: 0.5313 Epoch 2/15 - 92/92 [==============================] - 6s 64ms/step - loss: 1.0170 - accuracy: 0.5943 - val_loss: 0.9563 - val_accuracy: 0.6281 + 92/92 [==============================] - 6s 63ms/step - loss: 1.0828 - accuracy: 0.5746 - val_loss: 0.9974 - val_accuracy: 0.5981 Epoch 3/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.9168 - accuracy: 0.6468 - val_loss: 0.8525 - val_accuracy: 0.6553 + 92/92 [==============================] - 6s 63ms/step - loss: 0.9947 - accuracy: 0.6015 - val_loss: 0.9455 - val_accuracy: 0.6267 Epoch 4/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.8412 - accuracy: 0.6754 - val_loss: 0.9478 - val_accuracy: 0.6417 + 92/92 [==============================] - 6s 63ms/step - loss: 0.9154 - accuracy: 0.6482 - val_loss: 0.8459 - val_accuracy: 0.6771 Epoch 5/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.7881 - accuracy: 0.6856 - val_loss: 0.8132 - val_accuracy: 0.6839 + 92/92 [==============================] - 6s 63ms/step - loss: 0.8525 - accuracy: 0.6812 - val_loss: 0.8378 - val_accuracy: 0.6717 Epoch 6/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.7470 - accuracy: 0.7163 - val_loss: 0.8087 - val_accuracy: 0.6907 + 92/92 [==============================] - 6s 63ms/step - loss: 0.8104 - accuracy: 0.6948 - val_loss: 0.8545 - val_accuracy: 0.6567 Epoch 7/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.7111 - accuracy: 0.7302 - val_loss: 0.7582 - val_accuracy: 0.7234 + 92/92 [==============================] - 6s 63ms/step - loss: 0.7598 - accuracy: 0.6999 - val_loss: 0.8096 - val_accuracy: 0.6921 Epoch 8/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.6980 - accuracy: 0.7398 - val_loss: 0.7545 - val_accuracy: 0.7180 + 92/92 [==============================] - 6s 64ms/step - loss: 0.7397 - accuracy: 0.7166 - val_loss: 0.8358 - val_accuracy: 0.6812 Epoch 9/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.6437 - accuracy: 0.7578 - val_loss: 0.7517 - val_accuracy: 0.7003 + 92/92 [==============================] - 6s 64ms/step - loss: 0.7121 - accuracy: 0.7333 - val_loss: 0.7644 - val_accuracy: 0.6880 Epoch 10/15 - 92/92 [==============================] - 6s 63ms/step - loss: 0.6150 - accuracy: 0.7711 - val_loss: 0.7419 - val_accuracy: 0.7139 + 92/92 [==============================] - 6s 63ms/step - loss: 0.6739 - accuracy: 0.7449 - val_loss: 0.7528 - val_accuracy: 0.7084 Epoch 11/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.5926 - accuracy: 0.7715 - val_loss: 0.7543 - val_accuracy: 0.7248 + 92/92 [==============================] - 6s 63ms/step - loss: 0.6442 - accuracy: 0.7568 - val_loss: 0.7190 - val_accuracy: 0.7207 Epoch 12/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.5702 - accuracy: 0.7841 - val_loss: 0.6891 - val_accuracy: 0.7466 + 92/92 [==============================] - 6s 64ms/step - loss: 0.6113 - accuracy: 0.7715 - val_loss: 0.7588 - val_accuracy: 0.7057 Epoch 13/15 - 92/92 [==============================] - 6s 63ms/step - loss: 0.5467 - accuracy: 0.7977 - val_loss: 0.7306 - val_accuracy: 0.7234 + 92/92 [==============================] - 6s 63ms/step - loss: 0.5751 - accuracy: 0.7800 - val_loss: 0.7641 - val_accuracy: 0.7112 Epoch 14/15 - 92/92 [==============================] - 6s 64ms/step - loss: 0.5182 - accuracy: 0.8052 - val_loss: 0.7316 - val_accuracy: 0.7371 + 92/92 [==============================] - 6s 64ms/step - loss: 0.5595 - accuracy: 0.7847 - val_loss: 0.6969 - val_accuracy: 0.7357 Epoch 15/15 - 92/92 [==============================] - 6s 63ms/step - loss: 0.4997 - accuracy: 0.8048 - val_loss: 0.6754 - val_accuracy: 0.7411 + 92/92 [==============================] - 6s 63ms/step - loss: 0.5338 - accuracy: 0.8001 - val_loss: 0.7533 - val_accuracy: 0.7193 -Visualize Training Results --------------------------- +Visualize Training Results `⇑ <#top>`__ +############################################################################################################################### + After applying data augmentation and Dropout, there is less overfitting than before, and training and validation accuracy are closer aligned. @@ -758,8 +809,9 @@ than before, and training and validation accuracy are closer aligned. .. image:: 301-tensorflow-training-openvino-with-output_files/301-tensorflow-training-openvino-with-output_65_0.png -Predict on New Data -------------------- +Predict on New Data `⇑ <#top>`__ +############################################################################################################################### + Finally, let us use the model to classify an image that was not included in the training or validation sets. @@ -790,11 +842,12 @@ in the training or validation sets. .. parsed-literal:: 1/1 [==============================] - 0s 71ms/step - This image most likely belongs to sunflowers with a 97.75 percent confidence. + This image most likely belongs to sunflowers with a 88.60 percent confidence. -Save the TensorFlow Model -------------------------- +Save the TensorFlow Model `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -807,41 +860,41 @@ Save the TensorFlow Model .. parsed-literal:: - 2023-07-11 23:50:18.571028: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'random_flip_input' with dtype float and shape [?,180,180,3] + 2023-08-16 01:10:28.122100: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'random_flip_input' with dtype float and shape [?,180,180,3] [[{{node random_flip_input}}]] - 2023-07-11 23:50:18.656829: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-08-16 01:10:28.230661: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:50:18.666782: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'random_flip_input' with dtype float and shape [?,180,180,3] + 2023-08-16 01:10:28.240529: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'random_flip_input' with dtype float and shape [?,180,180,3] [[{{node random_flip_input}}]] - 2023-07-11 23:50:18.677553: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-08-16 01:10:28.251530: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:50:18.684368: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-08-16 01:10:28.258320: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:50:18.691156: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-08-16 01:10:28.265208: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:50:18.701873: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-08-16 01:10:28.275900: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:50:18.740670: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'sequential_1_input' with dtype float and shape [?,180,180,3] + 2023-08-16 01:10:28.314815: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'sequential_1_input' with dtype float and shape [?,180,180,3] [[{{node sequential_1_input}}]] - 2023-07-11 23:50:18.807441: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-08-16 01:10:28.381415: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:50:18.827744: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'sequential_1_input' with dtype float and shape [?,180,180,3] + 2023-08-16 01:10:28.401720: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'sequential_1_input' with dtype float and shape [?,180,180,3] [[{{node sequential_1_input}}]] - 2023-07-11 23:50:18.866385: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,22,22,64] + 2023-08-16 01:10:28.440601: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,22,22,64] [[{{node inputs}}]] - 2023-07-11 23:50:18.889771: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-08-16 01:10:28.464020: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:50:18.963275: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-08-16 01:10:28.537546: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:50:19.105456: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-08-16 01:10:28.678691: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:50:19.242700: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,22,22,64] + 2023-08-16 01:10:28.815557: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,22,22,64] [[{{node inputs}}]] - 2023-07-11 23:50:19.276405: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-08-16 01:10:28.849161: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:50:19.304418: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-08-16 01:10:28.877177: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] - 2023-07-11 23:50:19.350826: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] + 2023-08-16 01:10:28.923274: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'inputs' with dtype float and shape [?,180,180,3] [[{{node inputs}}]] WARNING:absl:Found untraced functions such as _jit_compiled_convolution_op, _jit_compiled_convolution_op, _jit_compiled_convolution_op, _update_step_xla while saving (showing 4 of 4). These functions will not be directly callable after loading. @@ -856,13 +909,12 @@ Save the TensorFlow Model INFO:tensorflow:Assets written to: model/flower/saved_model/assets -Convert the TensorFlow model with OpenVINO Model Optimizer ----------------------------------------------------------- +Convert the TensorFlow model with OpenVINO Model Optimizer `⇑ <#top>`__ +############################################################################################################################### -Use Model Optimizer Python API to convert the model to OpenVINO IR with -``FP16`` precision. For more information about Model Optimizer Python -API, see the `Model Optimizer Developer -Guide `__. +To convert the model to OpenVINO IR with ``FP16`` precision, use model +conversion Python API. For more information, see this +`page `__. .. code:: ipython3 @@ -872,8 +924,9 @@ Guide `__. ir_model = mo.convert_model(saved_model_dir=saved_model_dir, input_shape=[1,180,180,3], compress_to_fp16=True) serialize(ir_model, str(ir_model_path / "flower_ir.xml")) -Preprocessing Image Function ----------------------------- +Preprocessing Image Function `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -889,28 +942,55 @@ Preprocessing Image Function return input_image -OpenVINO Inference Engine Setup -------------------------------- +OpenVINO Runtime Setup `⇑ <#top>`__ +############################################################################################################################### + + +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + core = Core() + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + .. code:: ipython3 class_names=["daisy", "dandelion", "roses", "sunflowers", "tulips"] # Initialize OpenVINO runtime - ie = Core() - - # Neural Compute Stick - # compile the model for the CPU (you can choose manually CPU, GPU, etc.) - # or let the engine choose the best available device (AUTO) - compiled_model = ie.compile_model(model=ir_model, device_name="CPU") + core = Core() + compiled_model = core.compile_model(model=ir_model, device_name=device.value) del ir_model input_layer = compiled_model.input(0) output_layer = compiled_model.output(0) -Run the Inference Step ----------------------- +Run the Inference Step `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -948,21 +1028,22 @@ Run the Inference Step 'output/A_Close_Up_Photo_of_a_Dandelion.jpg' already exists. (1, 180, 180, 3) [1,180,180,3] - This image most likely belongs to dandelion with a 97.68 percent confidence. + This image most likely belongs to dandelion with a 98.50 percent confidence. -.. image:: 301-tensorflow-training-openvino-with-output_files/301-tensorflow-training-openvino-with-output_77_1.png +.. image:: 301-tensorflow-training-openvino-with-output_files/301-tensorflow-training-openvino-with-output_78_1.png -The Next Steps --------------- +The Next Steps `⇑ <#top>`__ +############################################################################################################################### + This tutorial showed how to train a TensorFlow model, how to convert that model to OpenVINO’s IR format, and how to do inference on the converted model. For faster inference speed, you can quantize the IR model. To see how to quantize this model with OpenVINO’s `Post-training Quantization with NNCF -Tool `__, +Tool `__, check out the `Post-Training Quantization with TensorFlow Classification Model <./301-tensorflow-training-openvino-nncf.ipynb>`__ notebook. diff --git a/docs/notebooks/301-tensorflow-training-openvino-with-output_files/301-tensorflow-training-openvino-with-output_56_1.png b/docs/notebooks/301-tensorflow-training-openvino-with-output_files/301-tensorflow-training-openvino-with-output_56_1.png index f49f6881f93..120bd04601b 100644 --- a/docs/notebooks/301-tensorflow-training-openvino-with-output_files/301-tensorflow-training-openvino-with-output_56_1.png +++ b/docs/notebooks/301-tensorflow-training-openvino-with-output_files/301-tensorflow-training-openvino-with-output_56_1.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:6fc6cb28f3e74c9f6e7e708ac55b401897025fb2c38aa974c5365167544f9890 -size 939080 +oid sha256:d3c1e81c182d56bc4a02c606b3599e6e93dce43deb38f42554c0b42d1db83bf6 +size 360658 diff --git a/docs/notebooks/301-tensorflow-training-openvino-with-output_files/301-tensorflow-training-openvino-with-output_65_0.png b/docs/notebooks/301-tensorflow-training-openvino-with-output_files/301-tensorflow-training-openvino-with-output_65_0.png index 8b0d8a3936a..4467d51e6ec 100644 --- a/docs/notebooks/301-tensorflow-training-openvino-with-output_files/301-tensorflow-training-openvino-with-output_65_0.png +++ b/docs/notebooks/301-tensorflow-training-openvino-with-output_files/301-tensorflow-training-openvino-with-output_65_0.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:000913f6826dabea02d8a1b6c692afe42402f7aa3a86889f6793f729151266a1 -size 56145 +oid sha256:54d4971bb60ae07f689a9f0a9c9938c0ba9ae7610cf75a03358e78e8d4afa002 +size 55759 diff --git a/docs/notebooks/301-tensorflow-training-openvino-with-output_files/301-tensorflow-training-openvino-with-output_77_1.png b/docs/notebooks/301-tensorflow-training-openvino-with-output_files/301-tensorflow-training-openvino-with-output_78_1.png similarity index 100% rename from docs/notebooks/301-tensorflow-training-openvino-with-output_files/301-tensorflow-training-openvino-with-output_77_1.png rename to docs/notebooks/301-tensorflow-training-openvino-with-output_files/301-tensorflow-training-openvino-with-output_78_1.png diff --git a/docs/notebooks/301-tensorflow-training-openvino-with-output_files/index.html b/docs/notebooks/301-tensorflow-training-openvino-with-output_files/index.html index 4f04a6182c4..e260abcfe3c 100644 --- a/docs/notebooks/301-tensorflow-training-openvino-with-output_files/index.html +++ b/docs/notebooks/301-tensorflow-training-openvino-with-output_files/index.html @@ -1,18 +1,18 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/301-tensorflow-training-openvino-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/301-tensorflow-training-openvino-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/301-tensorflow-training-openvino-with-output_files/


../
-301-tensorflow-training-openvino-with-output_13..> 12-Jul-2023 00:11                7042
-301-tensorflow-training-openvino-with-output_13..> 12-Jul-2023 00:11               64525
-301-tensorflow-training-openvino-with-output_14..> 12-Jul-2023 00:11               20653
-301-tensorflow-training-openvino-with-output_14..> 12-Jul-2023 00:11              167334
-301-tensorflow-training-openvino-with-output_16..> 12-Jul-2023 00:11               15872
-301-tensorflow-training-openvino-with-output_16..> 12-Jul-2023 00:11              225545
-301-tensorflow-training-openvino-with-output_17..> 12-Jul-2023 00:11               23154
-301-tensorflow-training-openvino-with-output_17..> 12-Jul-2023 00:11              154227
-301-tensorflow-training-openvino-with-output_28..> 12-Jul-2023 00:11              941151
-301-tensorflow-training-openvino-with-output_56..> 12-Jul-2023 00:11              939080
-301-tensorflow-training-openvino-with-output_65..> 12-Jul-2023 00:11               56145
-301-tensorflow-training-openvino-with-output_77..> 12-Jul-2023 00:11              143412
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/301-tensorflow-training-openvino-with-output_files/


../
+301-tensorflow-training-openvino-with-output_13..> 16-Aug-2023 01:31                7042
+301-tensorflow-training-openvino-with-output_13..> 16-Aug-2023 01:31               64525
+301-tensorflow-training-openvino-with-output_14..> 16-Aug-2023 01:31               20653
+301-tensorflow-training-openvino-with-output_14..> 16-Aug-2023 01:31              167334
+301-tensorflow-training-openvino-with-output_16..> 16-Aug-2023 01:31               15872
+301-tensorflow-training-openvino-with-output_16..> 16-Aug-2023 01:31              225545
+301-tensorflow-training-openvino-with-output_17..> 16-Aug-2023 01:31               23154
+301-tensorflow-training-openvino-with-output_17..> 16-Aug-2023 01:31              154227
+301-tensorflow-training-openvino-with-output_28..> 16-Aug-2023 01:31              941151
+301-tensorflow-training-openvino-with-output_56..> 16-Aug-2023 01:31              360658
+301-tensorflow-training-openvino-with-output_65..> 16-Aug-2023 01:31               55759
+301-tensorflow-training-openvino-with-output_78..> 16-Aug-2023 01:31              143412
 

diff --git a/docs/notebooks/302-pytorch-quantization-aware-training-with-output.rst b/docs/notebooks/302-pytorch-quantization-aware-training-with-output.rst index c84d2d8524d..693329f641d 100644 --- a/docs/notebooks/302-pytorch-quantization-aware-training-with-output.rst +++ b/docs/notebooks/302-pytorch-quantization-aware-training-with-output.rst @@ -1,6 +1,8 @@ Quantization Aware Training with NNCF, using PyTorch framework ============================================================== +.. _top: + This notebook is based on `ImageNet training in PyTorch `__. @@ -29,8 +31,25 @@ hub `__. **NOTE**: This notebook requires a C++ compiler. -Imports and Settings --------------------- +**Table of contents**: + +- `Imports and Settings <#imports-and-settings>`__ +- `Pre-train Floating-Point Model <#pre-train-floating-point-model>`__ + + - `Train Function <#train-function>`__ + - `Validate Function <#validate-function>`__ + - `Helpers <#helpers>`__ + - `Get a Pre-trained FP32 Model <#get-a-pre-trained-fp32-model>`__ + +- `Create and Initialize Quantization <#create-and-initialize-quantization>`__ +- `Fine-tune the Compressed Model <#fine-tune-the-compressed-model>`__ +- `Export INT8 Model to ONNX <#export-int8-model-to-onnx>`__ +- `Convert ONNX models to OpenVINO Intermediate Representation (IR) <#convert-onnx-models-to-openvino-intermediate-representation-ir>`__ +- `Benchmark Model Performance by Computing Inference Time <#benchmark-model-performance-by-computing-inference-time>`__ + +Imports and Settings `⇑ <#top>`__ +############################################################################################################################### + On Windows, add the required C++ directories to the system PATH. @@ -141,10 +160,10 @@ of the models will be stored. .. parsed-literal:: - 2023-07-11 23:50:27.854440: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. - 2023-07-11 23:50:27.889086: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + 2023-08-16 01:10:37.605341: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-08-16 01:10:37.639047: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. - 2023-07-11 23:50:28.441251: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + 2023-08-16 01:10:38.206632: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT .. parsed-literal:: @@ -172,14 +191,15 @@ of the models will be stored. .. parsed-literal:: - PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/302-pytorch-quantization-aware-training/model/resnet18_fp32.pth') + PosixPath('/opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/302-pytorch-quantization-aware-training/model/resnet18_fp32.pth') Download Tiny ImageNet dataset -* 100k images of shape 3x64x64 -* 200 different classes: snake, spider, cat, truck, grasshopper, gull, etc. +- 100k images of shape 3x64x64 +- 200 different classes: snake, spider, cat, truck, grasshopper, gull, + etc. .. code:: ipython3 @@ -230,21 +250,22 @@ Download Tiny ImageNet dataset Successfully downloaded and prepared dataset at: data/tiny-imagenet-200 -Pre-train Floating-Point Model ------------------------------- +Pre-train Floating-Point Model `⇑ <#top>`__ +############################################################################################################################### -Using NNCF for model compression assumes that a pre-trained model and a -training pipeline are already in use. +Using NNCF for model compression assumes that a pre-trained model and a training pipeline are +already in use. This tutorial demonstrates one possible training pipeline: a ResNet-18 model pre-trained on 1000 classes from ImageNet is fine-tuned with 200 -classes from Tiny-Imagenet. +classes from Tiny-ImageNet. Subsequently, the training and validation functions will be reused as is for quantization-aware training. -Train Function -~~~~~~~~~~~~~~ +Train Function `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -288,8 +309,9 @@ Train Function if i % print_frequency == 0: progress.display(i) -Validate Function -~~~~~~~~~~~~~~~~~ +Validate Function `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -330,8 +352,9 @@ Validate Function print(" * Acc@1 {top1.avg:.3f} Acc@5 {top5.avg:.3f}".format(top1=top1, top5=top5)) return top1.avg -Helpers -~~~~~~~ +Helpers `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -393,8 +416,9 @@ Helpers res.append(correct_k.mul_(100.0 / batch_size)) return res -Get a Pre-trained FP32 Model -~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Get a Pre-trained FP32 Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + А pre-trained floating-point model is a prerequisite for quantization. It can be obtained by tuning from scratch with the code below. However, @@ -460,9 +484,9 @@ section at the top of this notebook. .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torchvision/models/_utils.py:208: UserWarning: The parameter 'pretrained' is deprecated since 0.13 and may be removed in the future, please use 'weights' instead. + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torchvision/models/_utils.py:208: UserWarning: The parameter 'pretrained' is deprecated since 0.13 and may be removed in the future, please use 'weights' instead. warnings.warn( - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torchvision/models/_utils.py:223: UserWarning: Arguments other than a weight enum or `None` for 'weights' are deprecated since 0.13 and may be removed in the future. The current behavior is equivalent to passing `weights=None`. + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/torchvision/models/_utils.py:223: UserWarning: Arguments other than a weight enum or `None` for 'weights' are deprecated since 0.13 and may be removed in the future. The current behavior is equivalent to passing `weights=None`. warnings.warn(msg) @@ -519,8 +543,9 @@ Toolkit, to benchmark it in comparison with the ``INT8`` model. FP32 ONNX model was exported to output/resnet18_fp32.onnx. -Create and Initialize Quantization ----------------------------------- +Create and Initialize Quantization `⇑ <#top>`__ +############################################################################################################################### + NNCF enables compression-aware training by integrating into regular training pipelines. The framework is designed so that modifications to @@ -568,20 +593,21 @@ demonstrated here. .. parsed-literal:: - Test: [ 0/79] Time 0.194 (0.194) Loss 0.981 (0.981) Acc@1 78.91 (78.91) Acc@5 89.84 (89.84) - Test: [10/79] Time 0.151 (0.165) Loss 1.905 (1.623) Acc@1 46.88 (60.51) Acc@5 82.03 (84.09) - Test: [20/79] Time 0.150 (0.161) Loss 1.734 (1.692) Acc@1 63.28 (58.63) Acc@5 79.69 (83.04) - Test: [30/79] Time 0.153 (0.158) Loss 2.282 (1.781) Acc@1 50.00 (57.31) Acc@5 69.53 (81.50) - Test: [40/79] Time 0.150 (0.156) Loss 1.540 (1.825) Acc@1 62.50 (55.83) Acc@5 85.94 (80.96) - Test: [50/79] Time 0.161 (0.156) Loss 1.972 (1.820) Acc@1 57.03 (56.05) Acc@5 75.00 (80.73) - Test: [60/79] Time 0.151 (0.155) Loss 1.731 (1.846) Acc@1 57.81 (55.51) Acc@5 85.16 (80.21) - Test: [70/79] Time 0.151 (0.155) Loss 2.412 (1.872) Acc@1 47.66 (55.15) Acc@5 71.88 (79.61) + Test: [ 0/79] Time 0.161 (0.161) Loss 0.981 (0.981) Acc@1 78.91 (78.91) Acc@5 89.84 (89.84) + Test: [10/79] Time 0.145 (0.152) Loss 1.905 (1.623) Acc@1 46.88 (60.51) Acc@5 82.03 (84.09) + Test: [20/79] Time 0.149 (0.150) Loss 1.734 (1.692) Acc@1 63.28 (58.63) Acc@5 79.69 (83.04) + Test: [30/79] Time 0.148 (0.150) Loss 2.282 (1.781) Acc@1 50.00 (57.31) Acc@5 69.53 (81.50) + Test: [40/79] Time 0.148 (0.150) Loss 1.540 (1.825) Acc@1 62.50 (55.83) Acc@5 85.94 (80.96) + Test: [50/79] Time 0.146 (0.150) Loss 1.972 (1.820) Acc@1 57.03 (56.05) Acc@5 75.00 (80.73) + Test: [60/79] Time 0.147 (0.150) Loss 1.731 (1.846) Acc@1 57.81 (55.51) Acc@5 85.16 (80.21) + Test: [70/79] Time 0.151 (0.150) Loss 2.412 (1.872) Acc@1 47.66 (55.15) Acc@5 71.88 (79.61) * Acc@1 55.540 Acc@5 80.200 Accuracy of initialized INT8 model: 55.540 -Fine-tune the Compressed Model ------------------------------- +Fine-tune the Compressed Model `⇑ <#top>`__ +############################################################################################################################### + At this step, a regular fine-tuning process is applied to further improve quantized model accuracy. Normally, several epochs of tuning are @@ -606,37 +632,38 @@ training pipeline are required. Here is a simple example. .. parsed-literal:: - Epoch:[0][ 0/782] Time 0.405 (0.405) Loss 0.740 (0.740) Acc@1 84.38 (84.38) Acc@5 96.88 (96.88) - Epoch:[0][ 50/782] Time 0.395 (0.389) Loss 0.911 (0.802) Acc@1 78.91 (80.15) Acc@5 92.97 (94.42) - Epoch:[0][100/782] Time 0.411 (0.389) Loss 0.631 (0.798) Acc@1 84.38 (80.24) Acc@5 95.31 (94.38) - Epoch:[0][150/782] Time 0.415 (0.388) Loss 0.836 (0.792) Acc@1 80.47 (80.48) Acc@5 94.53 (94.43) - Epoch:[0][200/782] Time 0.379 (0.388) Loss 0.873 (0.780) Acc@1 75.00 (80.65) Acc@5 94.53 (94.59) - Epoch:[0][250/782] Time 0.408 (0.388) Loss 0.735 (0.778) Acc@1 84.38 (80.77) Acc@5 95.31 (94.53) - Epoch:[0][300/782] Time 0.404 (0.388) Loss 0.615 (0.771) Acc@1 85.16 (80.99) Acc@5 97.66 (94.58) - Epoch:[0][350/782] Time 0.363 (0.388) Loss 0.599 (0.767) Acc@1 85.16 (81.14) Acc@5 95.31 (94.58) - Epoch:[0][400/782] Time 0.394 (0.388) Loss 0.798 (0.765) Acc@1 82.03 (81.21) Acc@5 92.97 (94.56) - Epoch:[0][450/782] Time 0.429 (0.388) Loss 0.630 (0.762) Acc@1 85.16 (81.26) Acc@5 96.88 (94.58) - Epoch:[0][500/782] Time 0.365 (0.389) Loss 0.633 (0.757) Acc@1 85.94 (81.45) Acc@5 96.88 (94.63) - Epoch:[0][550/782] Time 0.368 (0.388) Loss 0.749 (0.755) Acc@1 82.03 (81.49) Acc@5 92.97 (94.65) - Epoch:[0][600/782] Time 0.366 (0.389) Loss 0.927 (0.753) Acc@1 78.12 (81.53) Acc@5 88.28 (94.67) - Epoch:[0][650/782] Time 0.378 (0.389) Loss 0.645 (0.749) Acc@1 84.38 (81.60) Acc@5 95.31 (94.71) - Epoch:[0][700/782] Time 0.401 (0.389) Loss 0.816 (0.749) Acc@1 82.03 (81.62) Acc@5 91.41 (94.69) - Epoch:[0][750/782] Time 0.398 (0.389) Loss 0.811 (0.746) Acc@1 80.47 (81.69) Acc@5 94.53 (94.72) - Test: [ 0/79] Time 0.178 (0.178) Loss 1.092 (1.092) Acc@1 75.00 (75.00) Acc@5 86.72 (86.72) - Test: [10/79] Time 0.137 (0.141) Loss 1.917 (1.526) Acc@1 48.44 (62.64) Acc@5 78.12 (83.88) - Test: [20/79] Time 0.139 (0.139) Loss 1.631 (1.602) Acc@1 64.06 (60.68) Acc@5 81.25 (83.71) - Test: [30/79] Time 0.137 (0.139) Loss 2.037 (1.691) Acc@1 57.81 (59.25) Acc@5 71.09 (82.23) - Test: [40/79] Time 0.137 (0.138) Loss 1.563 (1.743) Acc@1 64.84 (58.02) Acc@5 82.81 (81.33) - Test: [50/79] Time 0.139 (0.138) Loss 1.926 (1.750) Acc@1 52.34 (57.77) Acc@5 76.56 (81.04) - Test: [60/79] Time 0.136 (0.138) Loss 1.559 (1.781) Acc@1 67.19 (57.24) Acc@5 84.38 (80.58) - Test: [70/79] Time 0.136 (0.137) Loss 2.353 (1.806) Acc@1 46.88 (56.81) Acc@5 72.66 (80.08) + Epoch:[0][ 0/782] Time 0.391 (0.391) Loss 0.740 (0.740) Acc@1 84.38 (84.38) Acc@5 96.88 (96.88) + Epoch:[0][ 50/782] Time 0.387 (0.383) Loss 0.911 (0.802) Acc@1 78.91 (80.15) Acc@5 92.97 (94.42) + Epoch:[0][100/782] Time 0.387 (0.384) Loss 0.631 (0.798) Acc@1 84.38 (80.24) Acc@5 95.31 (94.38) + Epoch:[0][150/782] Time 0.377 (0.383) Loss 0.836 (0.792) Acc@1 80.47 (80.48) Acc@5 94.53 (94.43) + Epoch:[0][200/782] Time 0.431 (0.385) Loss 0.873 (0.780) Acc@1 75.00 (80.65) Acc@5 94.53 (94.59) + Epoch:[0][250/782] Time 0.385 (0.386) Loss 0.735 (0.778) Acc@1 84.38 (80.77) Acc@5 95.31 (94.53) + Epoch:[0][300/782] Time 0.411 (0.386) Loss 0.615 (0.771) Acc@1 85.16 (80.99) Acc@5 97.66 (94.58) + Epoch:[0][350/782] Time 0.386 (0.386) Loss 0.599 (0.767) Acc@1 85.16 (81.14) Acc@5 95.31 (94.58) + Epoch:[0][400/782] Time 0.385 (0.386) Loss 0.798 (0.765) Acc@1 82.03 (81.21) Acc@5 92.97 (94.56) + Epoch:[0][450/782] Time 0.432 (0.386) Loss 0.630 (0.762) Acc@1 85.16 (81.26) Acc@5 96.88 (94.58) + Epoch:[0][500/782] Time 0.397 (0.386) Loss 0.633 (0.757) Acc@1 85.94 (81.45) Acc@5 96.88 (94.63) + Epoch:[0][550/782] Time 0.383 (0.387) Loss 0.749 (0.755) Acc@1 82.03 (81.49) Acc@5 92.97 (94.65) + Epoch:[0][600/782] Time 0.394 (0.387) Loss 0.927 (0.753) Acc@1 78.12 (81.53) Acc@5 88.28 (94.67) + Epoch:[0][650/782] Time 0.384 (0.387) Loss 0.645 (0.749) Acc@1 84.38 (81.60) Acc@5 95.31 (94.71) + Epoch:[0][700/782] Time 0.383 (0.387) Loss 0.816 (0.749) Acc@1 82.03 (81.62) Acc@5 91.41 (94.69) + Epoch:[0][750/782] Time 0.385 (0.387) Loss 0.811 (0.746) Acc@1 80.47 (81.69) Acc@5 94.53 (94.72) + Test: [ 0/79] Time 0.189 (0.189) Loss 1.092 (1.092) Acc@1 75.00 (75.00) Acc@5 86.72 (86.72) + Test: [10/79] Time 0.145 (0.154) Loss 1.917 (1.526) Acc@1 48.44 (62.64) Acc@5 78.12 (83.88) + Test: [20/79] Time 0.144 (0.149) Loss 1.631 (1.602) Acc@1 64.06 (60.68) Acc@5 81.25 (83.71) + Test: [30/79] Time 0.145 (0.148) Loss 2.037 (1.691) Acc@1 57.81 (59.25) Acc@5 71.09 (82.23) + Test: [40/79] Time 0.144 (0.147) Loss 1.563 (1.743) Acc@1 64.84 (58.02) Acc@5 82.81 (81.33) + Test: [50/79] Time 0.146 (0.147) Loss 1.926 (1.750) Acc@1 52.34 (57.77) Acc@5 76.56 (81.04) + Test: [60/79] Time 0.144 (0.146) Loss 1.559 (1.781) Acc@1 67.19 (57.24) Acc@5 84.38 (80.58) + Test: [70/79] Time 0.144 (0.146) Loss 2.353 (1.806) Acc@1 46.88 (56.81) Acc@5 72.66 (80.08) * Acc@1 57.320 Acc@5 80.730 Accuracy of tuned INT8 model: 57.320 Accuracy drop of tuned INT8 model over pre-trained FP32 model: -1.800 -Export INT8 Model to ONNX -------------------------- +Export INT8 Model to ONNX `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -650,7 +677,7 @@ Export INT8 Model to ONNX .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/nncf/torch/quantization/quantize_functions.py:140: FutureWarning: 'torch.onnx._patch_torch._graph_op' is deprecated in version 1.13 and will be removed in version 1.14. Please note 'g.op()' is to be removed from torch.Graph. Please open a GitHub issue if you need this functionality.. + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/nncf/torch/quantization/quantize_functions.py:140: FutureWarning: 'torch.onnx._patch_torch._graph_op' is deprecated in version 1.13 and will be removed in version 1.14. Please note 'g.op()' is to be removed from torch.Graph. Please open a GitHub issue if you need this functionality.. output = g.op( @@ -659,18 +686,17 @@ Export INT8 Model to ONNX INT8 ONNX model exported to output/resnet18_int8.onnx. -Convert ONNX models to OpenVINO Intermediate Representation (IR) ----------------------------------------------------------------- +Convert ONNX models to OpenVINO Intermediate Representation (IR). `⇑ <#top>`__ +############################################################################################################################### -Use Model Optimizer Python API to convert the ONNX model to OpenVINO IR, -with ``FP16`` precision. Then, add the mean values to the model and +Use model conversion Python API to convert the ONNX model to OpenVINO +IR, with ``FP16`` precision. Then, add the mean values to the model and scale the input with the standard deviation by the ``mean_values`` and ``scale_values`` parameters. It is not necessary to normalize input data before propagating it through the network with these options. -For more information about Model Optimizer Python API, see the `Model -Optimizer Developer -Guide `__. +For more information about model conversion, see this +`page `__. .. code:: ipython3 @@ -694,8 +720,9 @@ Guide `__. ) serialize(model, str(int8_ir_path)) -Benchmark Model Performance by Computing Inference Time -------------------------------------------------------- +Benchmark Model Performance by Computing Inference Time `⇑ <#top>`__ +############################################################################################################################### + Finally, measure the inference performance of the ``FP32`` and ``INT8`` models, using `Benchmark @@ -733,9 +760,9 @@ throughput (frames per second) values. .. parsed-literal:: Benchmark FP32 model (IR) - [ INFO ] Throughput: 2916.72 FPS + [ INFO ] Throughput: 2896.36 FPS Benchmark INT8 model (IR) - [ INFO ] Throughput: 11947.88 FPS + [ INFO ] Throughput: 12326.44 FPS Show CPU Information for reference. diff --git a/docs/notebooks/305-tensorflow-quantization-aware-training-with-output.rst b/docs/notebooks/305-tensorflow-quantization-aware-training-with-output.rst index 005c2e9fe4c..85acba00ec7 100644 --- a/docs/notebooks/305-tensorflow-quantization-aware-training-with-output.rst +++ b/docs/notebooks/305-tensorflow-quantization-aware-training-with-output.rst @@ -1,32 +1,45 @@ Quantization Aware Training with NNCF, using TensorFlow Framework ================================================================= +.. _top: + The goal of this notebook to demonstrate how to use the Neural Network Compression Framework `NNCF `__ 8-bit quantization to optimize a TensorFlow model for inference with OpenVINO™ Toolkit. The optimization process contains the following steps: -* Transforming the original ``FP32`` model to ``INT8``. -* Using fine-tuning to restore the accuracy. -* Exporting optimized and original models to Frozen Graph and then to OpenVINO. -* Measuring and comparing the performance of models. +- Transforming the original ``FP32`` model to ``INT8`` +- Using fine-tuning to restore the accuracy. +- Exporting optimized and original models to Frozen Graph and then to + OpenVINO. +- Measuring and comparing the performance of models. For more advanced usage, refer to these `examples `__. This tutorial uses the ResNet-18 model with Imagenette dataset. -Imagenette is a subset of 10 easily classified classes from the Imagenet +Imagenette is a subset of 10 easily classified classes from the ImageNet dataset. Using the smaller model and dataset will speed up training and download time. -Imports and Settings --------------------- +**Table of contents**: -Import NNCF and all auxiliary packages from your Python code. Set a name -for the model, input image size, used batch size, and the learning rate. -Also, define paths where Frozen Graph and OpenVINO IR versions of the -models will be stored. +- `Imports and Settings <#imports-and-settings>`__ +- `Dataset Preprocessing <#dataset-preprocessing>`__ +- `Define a Floating-Point Model <#define-a-floating-point-model>`__ +- `Pre-train a Floating-Point Model <#pre-train-a-floating-point-model>`__ +- `Create and Initialize Quantization <#create-and-initialize-quantization>`__ +- `Fine-tune the Compressed Model <#fine-tune-the-compressed-model>`__ +- `Export Models to OpenVINO Intermediate Representation (IR) <#export-models-to-openvino-intermediate-representation-ir>`__ +- `Benchmark Model Performance by Computing Inference Time <#benchmark-model-performance-by-computing-inference-time>`__ + +Imports and Settings `⇑ <#top>`__ +############################################################################################################################### + +Import NNCF and all auxiliary packages from your Python code. Set a name for the model, input image +size, used batch size, and the learning rate. Also, define paths where +Frozen Graph and OpenVINO IR versions of the models will be stored. **NOTE**: All NNCF logging messages below ERROR level (INFO and WARNING) are disabled to simplify the tutorial. For production use, @@ -35,9 +48,18 @@ models will be stored. .. code:: ipython3 - !pip install -q 'openvino-dev>=2023.0.0' 'nncf>=2.5.0' + !pip install -q "openvino-dev>=2023.0.0" "nncf>=2.5.0" !pip install -q "tensorflow-datasets>=4.8.0" + +.. parsed-literal:: + + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + ERROR: pip's dependency resolver does not currently take into account all the packages that are installed. This behaviour is the source of the following dependency conflicts. + pytorch-lightning 1.6.5 requires protobuf<=3.20.1, but you have protobuf 3.20.3 which is incompatible. + + .. code:: ipython3 from pathlib import Path @@ -85,10 +107,10 @@ models will be stored. .. parsed-literal:: - 2023-07-11 23:57:23.400449: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. - 2023-07-11 23:57:23.436263: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. + 2023-08-16 01:17:34.103410: I tensorflow/core/util/port.cc:110] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`. + 2023-08-16 01:17:34.137361: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. To enable the following instructions: AVX2 AVX512F AVX512_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. - 2023-07-11 23:57:24.023497: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT + 2023-08-16 01:17:34.726614: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT .. parsed-literal:: @@ -98,15 +120,22 @@ models will be stored. Downloading data from https://storage.openvinotoolkit.org/repositories/nncf/openvino_notebook_ckpts/305_resnet18_imagenette_fp32_v1.h5 134604992/134604992 [==============================] - 30s 0us/step Absolute path where the model weights are saved: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/305-tensorflow-quantization-aware-training/model/ResNet-18_fp32.h5 + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/305-tensorflow-quantization-aware-training/model/ResNet-18_fp32.h5 -Dataset Preprocessing ---------------------- +Dataset Preprocessing `⇑ <#top>`__ +############################################################################################################################### + Download and prepare Imagenette 160px dataset. - Number of classes: 10 - -Download size: 94.18 MiB \| Split \| Examples \| \|————–|———-\| \| -‘train’ \| 12,894 \| \| ‘validation’ \| 500 \| +Download size: 94.18 MiB + +:: + + | Split | Examples | + |--------------|----------| + | 'train' | 12,894 | + | 'validation' | 500 | .. code:: ipython3 @@ -118,17 +147,17 @@ Download size: 94.18 MiB \| Split \| Examples \| \|————–|———-\| .. parsed-literal:: - 2023-07-11 23:57:56.999149: W tensorflow/core/common_runtime/gpu/gpu_device.cc:1956] Cannot dlopen some GPU libraries. Please make sure the missing libraries mentioned above are installed properly if you would like to use GPU. Follow the guide at https://www.tensorflow.org/install/gpu for how to download and setup the required libraries for your platform. + 2023-08-16 01:18:08.016585: W tensorflow/core/common_runtime/gpu/gpu_device.cc:1956] Cannot dlopen some GPU libraries. Please make sure the missing libraries mentioned above are installed properly if you would like to use GPU. Follow the guide at https://www.tensorflow.org/install/gpu for how to download and setup the required libraries for your platform. Skipping registering GPU devices... - 2023-07-11 23:57:57.107813: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_3' with dtype int64 and shape [1] - [[{{node Placeholder/_3}}]] - 2023-07-11 23:57:57.108141: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [1] + 2023-08-16 01:18:08.132762: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_1' with dtype string and shape [1] + [[{{node Placeholder/_1}}]] + 2023-08-16 01:18:08.133087: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [1] [[{{node Placeholder/_0}}]] - 2023-07-11 23:57:57.153228: W tensorflow/core/kernels/data/cache_dataset_ops.cc:856] The calling iterator did not fully read the dataset being cached. In order to avoid unexpected truncation of the dataset, the partially cached contents of the dataset will be discarded. This can happen if you have an input pipeline similar to `dataset.cache().take(k).repeat()`. You should use `dataset.take(k).cache().repeat()` instead. + 2023-08-16 01:18:08.170026: W tensorflow/core/kernels/data/cache_dataset_ops.cc:856] The calling iterator did not fully read the dataset being cached. In order to avoid unexpected truncation of the dataset, the partially cached contents of the dataset will be discarded. This can happen if you have an input pipeline similar to `dataset.cache().take(k).repeat()`. You should use `dataset.take(k).cache().repeat()` instead. -.. image:: 305-tensorflow-quantization-aware-training-with-output_files/305-tensorflow-quantization-aware-training-with-output_5_1.png +.. image:: 305-tensorflow-quantization-aware-training-with-output_files/305-tensorflow-quantization-aware-training-with-output_6_1.png .. code:: ipython3 @@ -149,8 +178,9 @@ Download size: 94.18 MiB \| Split \| Examples \| \|————–|———-\| .batch(BATCH_SIZE) .prefetch(tf.data.experimental.AUTOTUNE)) -Define a Floating-Point Model ------------------------------ +Define a Floating-Point Model `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -224,8 +254,9 @@ Define a Floating-Point Model IMG_SHAPE = IMG_SIZE + (3,) fp32_model = ResNet18(input_shape=IMG_SHAPE) -Pre-train a Floating-Point Model --------------------------------- +Pre-train a Floating-Point Model `⇑ <#top>`__ +############################################################################################################################### + Using NNCF for model compression assumes that the user has a pre-trained model and a training pipeline. @@ -255,21 +286,22 @@ model and a training pipeline. .. parsed-literal:: - 2023-07-11 23:57:57.994768: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_2' with dtype string and shape [1] - [[{{node Placeholder/_2}}]] - 2023-07-11 23:57:57.995164: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int64 and shape [1] - [[{{node Placeholder/_4}}]] + 2023-08-16 01:18:09.025847: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_1' with dtype string and shape [1] + [[{{node Placeholder/_1}}]] + 2023-08-16 01:18:09.026203: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_0' with dtype string and shape [1] + [[{{node Placeholder/_0}}]] .. parsed-literal:: - 4/4 [==============================] - 1s 235ms/sample - loss: 0.9807 - acc@1: 0.8220 + 4/4 [==============================] - 1s 229ms/sample - loss: 0.9807 - acc@1: 0.8220 Accuracy of FP32 model: 0.822 -Create and Initialize Quantization ----------------------------------- +Create and Initialize Quantization `⇑ <#top>`__ +############################################################################################################################### + NNCF enables compression-aware training by integrating into regular training pipelines. The framework is designed so that modifications to @@ -309,13 +341,13 @@ scenario and requires only 3 modifications. .. parsed-literal:: - 2023-07-11 23:58:00.692522: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_1' with dtype string and shape [1] - [[{{node Placeholder/_1}}]] - 2023-07-11 23:58:00.692903: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_2' with dtype string and shape [1] - [[{{node Placeholder/_2}}]] - 2023-07-11 23:58:01.596992: W tensorflow/core/kernels/data/cache_dataset_ops.cc:856] The calling iterator did not fully read the dataset being cached. In order to avoid unexpected truncation of the dataset, the partially cached contents of the dataset will be discarded. This can happen if you have an input pipeline similar to `dataset.cache().take(k).repeat()`. You should use `dataset.take(k).cache().repeat()` instead. - 2023-07-11 23:58:02.209552: W tensorflow/core/kernels/data/cache_dataset_ops.cc:856] The calling iterator did not fully read the dataset being cached. In order to avoid unexpected truncation of the dataset, the partially cached contents of the dataset will be discarded. This can happen if you have an input pipeline similar to `dataset.cache().take(k).repeat()`. You should use `dataset.take(k).cache().repeat()` instead. - 2023-07-11 23:58:10.535691: W tensorflow/core/kernels/data/cache_dataset_ops.cc:856] The calling iterator did not fully read the dataset being cached. In order to avoid unexpected truncation of the dataset, the partially cached contents of the dataset will be discarded. This can happen if you have an input pipeline similar to `dataset.cache().take(k).repeat()`. You should use `dataset.take(k).cache().repeat()` instead. + 2023-08-16 01:18:11.729441: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_4' with dtype int64 and shape [1] + [[{{node Placeholder/_4}}]] + 2023-08-16 01:18:11.729828: I tensorflow/core/common_runtime/executor.cc:1197] [/device:CPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): INVALID_ARGUMENT: You must feed a value for placeholder tensor 'Placeholder/_3' with dtype int64 and shape [1] + [[{{node Placeholder/_3}}]] + 2023-08-16 01:18:12.738622: W tensorflow/core/kernels/data/cache_dataset_ops.cc:856] The calling iterator did not fully read the dataset being cached. In order to avoid unexpected truncation of the dataset, the partially cached contents of the dataset will be discarded. This can happen if you have an input pipeline similar to `dataset.cache().take(k).repeat()`. You should use `dataset.take(k).cache().repeat()` instead. + 2023-08-16 01:18:13.389616: W tensorflow/core/kernels/data/cache_dataset_ops.cc:856] The calling iterator did not fully read the dataset being cached. In order to avoid unexpected truncation of the dataset, the partially cached contents of the dataset will be discarded. This can happen if you have an input pipeline similar to `dataset.cache().take(k).repeat()`. You should use `dataset.take(k).cache().repeat()` instead. + 2023-08-16 01:18:21.360841: W tensorflow/core/kernels/data/cache_dataset_ops.cc:856] The calling iterator did not fully read the dataset being cached. In order to avoid unexpected truncation of the dataset, the partially cached contents of the dataset will be discarded. This can happen if you have an input pipeline similar to `dataset.cache().take(k).repeat()`. You should use `dataset.take(k).cache().repeat()` instead. Evaluate the new model on the validation set after initialization of @@ -341,11 +373,12 @@ demonstrated here. .. parsed-literal:: - 4/4 [==============================] - 1s 300ms/sample - loss: 0.9766 - acc@1: 0.8120 + 4/4 [==============================] - 1s 301ms/sample - loss: 0.9766 - acc@1: 0.8120 -Fine-tune the Compressed Model ------------------------------- +Fine-tune the Compressed Model `⇑ <#top>`__ +############################################################################################################################### + At this step, a regular fine-tuning process is applied to further improve quantized model accuracy. Normally, several epochs of tuning are @@ -373,24 +406,24 @@ training pipeline are required. Here is a simple example. Accuracy of INT8 model after initialization: 0.812 Epoch 1/2 - 101/101 [==============================] - 48s 417ms/step - loss: 0.7134 - acc@1: 0.9299 + 101/101 [==============================] - 49s 417ms/step - loss: 0.7134 - acc@1: 0.9299 Epoch 2/2 - 101/101 [==============================] - 42s 419ms/step - loss: 0.6807 - acc@1: 0.9489 - 4/4 [==============================] - 1s 141ms/sample - loss: 0.9760 - acc@1: 0.8160 + 101/101 [==============================] - 42s 414ms/step - loss: 0.6807 - acc@1: 0.9489 + 4/4 [==============================] - 1s 144ms/sample - loss: 0.9760 - acc@1: 0.8160 Accuracy of INT8 model after fine-tuning: 0.816 Accuracy drop of tuned INT8 model over pre-trained FP32 model: 0.006 -Export Models to OpenVINO Intermediate Representation (IR) ----------------------------------------------------------- +Export Models to OpenVINO Intermediate Representation (IR) `⇑ <#top>`__ +############################################################################################################################### -Use Model Optimizer Python API to convert the models to OpenVINO IR. -For more information about Model Optimizer, see the `Model Optimizer -Developer -Guide `__. +Use model conversion Python API to convert the models to OpenVINO IR. + +For more information about model conversion, see this +`page `__. Executing this command may take a while. @@ -404,9 +437,9 @@ Executing this command may take a while. .. parsed-literal:: - 2023-07-11 23:59:43.746206: I tensorflow/core/grappler/devices.cc:66] Number of eligible GPUs (core count >= 8, compute capability >= 0.0): 2 - 2023-07-11 23:59:43.746302: I tensorflow/core/grappler/clusters/single_machine.cc:358] Starting new session - 2023-07-11 23:59:43.881725: W tensorflow/core/common_runtime/gpu/gpu_device.cc:1956] Cannot dlopen some GPU libraries. Please make sure the missing libraries mentioned above are installed properly if you would like to use GPU. Follow the guide at https://www.tensorflow.org/install/gpu for how to download and setup the required libraries for your platform. + 2023-08-16 01:19:54.530759: I tensorflow/core/grappler/devices.cc:66] Number of eligible GPUs (core count >= 8, compute capability >= 0.0): 2 + 2023-08-16 01:19:54.530838: I tensorflow/core/grappler/clusters/single_machine.cc:358] Starting new session + 2023-08-16 01:19:54.651453: W tensorflow/core/common_runtime/gpu/gpu_device.cc:1956] Cannot dlopen some GPU libraries. Please make sure the missing libraries mentioned above are installed properly if you would like to use GPU. Follow the guide at https://www.tensorflow.org/install/gpu for how to download and setup the required libraries for your platform. Skipping registering GPU devices... @@ -420,14 +453,15 @@ Executing this command may take a while. .. parsed-literal:: - 2023-07-11 23:59:45.484067: I tensorflow/core/grappler/devices.cc:66] Number of eligible GPUs (core count >= 8, compute capability >= 0.0): 2 - 2023-07-11 23:59:45.484141: I tensorflow/core/grappler/clusters/single_machine.cc:358] Starting new session - 2023-07-11 23:59:45.485613: W tensorflow/core/common_runtime/gpu/gpu_device.cc:1956] Cannot dlopen some GPU libraries. Please make sure the missing libraries mentioned above are installed properly if you would like to use GPU. Follow the guide at https://www.tensorflow.org/install/gpu for how to download and setup the required libraries for your platform. + 2023-08-16 01:19:56.200644: I tensorflow/core/grappler/devices.cc:66] Number of eligible GPUs (core count >= 8, compute capability >= 0.0): 2 + 2023-08-16 01:19:56.200714: I tensorflow/core/grappler/clusters/single_machine.cc:358] Starting new session + 2023-08-16 01:19:56.202200: W tensorflow/core/common_runtime/gpu/gpu_device.cc:1956] Cannot dlopen some GPU libraries. Please make sure the missing libraries mentioned above are installed properly if you would like to use GPU. Follow the guide at https://www.tensorflow.org/install/gpu for how to download and setup the required libraries for your platform. Skipping registering GPU devices... -Benchmark Model Performance by Computing Inference Time -------------------------------------------------------- +Benchmark Model Performance by Computing Inference Time `⇑ <#top>`__ +############################################################################################################################### + Finally, measure the inference performance of the ``FP32`` and ``INT8`` models, using `Benchmark @@ -469,10 +503,10 @@ throughput (frames per second) values. .. parsed-literal:: Benchmark FP32 model (IR) - [ INFO ] Throughput: 2843.91 FPS + [ INFO ] Throughput: 2831.57 FPS Benchmark INT8 model (IR) - [ INFO ] Throughput: 11931.91 FPS + [ INFO ] Throughput: 11769.65 FPS Show CPU Information for reference. diff --git a/docs/notebooks/305-tensorflow-quantization-aware-training-with-output_files/305-tensorflow-quantization-aware-training-with-output_5_1.png b/docs/notebooks/305-tensorflow-quantization-aware-training-with-output_files/305-tensorflow-quantization-aware-training-with-output_6_1.png similarity index 100% rename from docs/notebooks/305-tensorflow-quantization-aware-training-with-output_files/305-tensorflow-quantization-aware-training-with-output_5_1.png rename to docs/notebooks/305-tensorflow-quantization-aware-training-with-output_files/305-tensorflow-quantization-aware-training-with-output_6_1.png diff --git a/docs/notebooks/305-tensorflow-quantization-aware-training-with-output_files/index.html b/docs/notebooks/305-tensorflow-quantization-aware-training-with-output_files/index.html index 66d6e56fc17..015ff50edb4 100644 --- a/docs/notebooks/305-tensorflow-quantization-aware-training-with-output_files/index.html +++ b/docs/notebooks/305-tensorflow-quantization-aware-training-with-output_files/index.html @@ -1,7 +1,7 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/305-tensorflow-quantization-aware-training-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/305-tensorflow-quantization-aware-training-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/305-tensorflow-quantization-aware-training-with-output_files/


../
-305-tensorflow-quantization-aware-training-with..> 12-Jul-2023 00:11              519560
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/305-tensorflow-quantization-aware-training-with-output_files/


../
+305-tensorflow-quantization-aware-training-with..> 16-Aug-2023 01:31              519560
 

diff --git a/docs/notebooks/401-object-detection-with-output.rst b/docs/notebooks/401-object-detection-with-output.rst index f2e12651410..bc83f4a2af3 100644 --- a/docs/notebooks/401-object-detection-with-output.rst +++ b/docs/notebooks/401-object-detection-with-output.rst @@ -1,6 +1,8 @@ Live Object Detection with OpenVINO™ ==================================== +.. _top: + This notebook demonstrates live object detection with OpenVINO, using the `SSDLite MobileNetV2 `__ @@ -9,16 +11,44 @@ Zoo `__. Final part of this notebook shows live inference results from a webcam. Additionally, you can also upload a video file. - **NOTE**: To use this notebook with a webcam, you need to run the - notebook on a computer with a webcam. If you run the notebook on a - server, the webcam will not work. However, you can still do inference - on a video. +.. note:: -Preparation ------------ + To use this notebook with a webcam, you need to run the notebook on a computer + with a webcam. If you run the notebook on a server, the webcam will not work. + However, you can still do inference on a video. + +**Table of contents**: + +- `Preparation <#preparation>`__ + + - `Install requirements <#install-requirements>`__ + - `Imports <#imports>`__ + +- `The Model <#the-model>`__ + + - `Download the Model <#download-the-model>`__ + - `Convert the Model <#convert-the-model>`__ + - `Load the Model <#load-the-model>`__ + +- `Processing <#processing>`__ + + - `Process Results <#process-results>`__ + - `Main Processing Function <#main-processing-function>`__ + +- `Run <#run>`__ + + - `Run Live Object Detection <#run-live-object-detection>`__ + - `Run Object Detection on a Video File <#run-object-detection-on-a-video-file>`__ + +- `References <#references>`__ + +Preparation `⇑ <#top>`__ +############################################################################################################################### + + +Install requirements `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Install requirements -~~~~~~~~~~~~~~~~~~~~ .. code:: ipython3 @@ -34,16 +64,24 @@ Install requirements ) +.. parsed-literal:: + + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + + .. parsed-literal:: - ('notebook_utils.py', ) + ('notebook_utils.py', ) -Imports -~~~~~~~ +Imports `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -61,11 +99,13 @@ Imports import notebook_utils as utils -The Model ---------- +The Model `⇑ <#top>`__ +############################################################################################################################### + + +Download the Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Download the Model -~~~~~~~~~~~~~~~~~~ Use the ``download_file``, a function from the ``notebook_utils`` file. It automatically creates a directory structure and downloads the @@ -74,9 +114,10 @@ downloaded and unpacked. The chosen model comes from the public directory, which means it must be converted into OpenVINO Intermediate Representation (OpenVINO IR). - **NOTE**: Using a model other than ``ssdlite_mobilenet_v2`` may - require different conversion parameters as well as pre- and - post-processing. +.. note:: + + Using a model other than ``ssdlite_mobilenet_v2`` may require different + conversion parameters as well as pre- and post-processing. .. code:: ipython3 @@ -107,12 +148,13 @@ Representation (OpenVINO IR). model/ssdlite_mobilenet_v2_coco_2018_05_09.tar.gz: 0%| | 0.00/48.7M [00:00`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The pre-trained model is in TensorFlow format. To use it with OpenVINO, -convert it to OpenVINO IR format using `Model Optimizer Python -API `__ +convert it to OpenVINO IR format, using `model conversion Python +API `__ (``mo.convert_model`` function). If the model has been already converted, this step is skipped. @@ -141,8 +183,9 @@ converted, this step is skipped. [ WARNING ] The Preprocessor block has been removed. Only nodes performing mean value subtraction and scaling (if applicable) are kept. -Load the Model -~~~~~~~~~~~~~~ +Load the Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Only a few lines of code are required to run the model. First, initialize OpenVINO Runtime. Then, read the network architecture and @@ -151,19 +194,41 @@ desired device. If you choose ``GPU`` you need to wait for a while, as the startup time is much longer than in the case of ``CPU``. There is a possibility to let OpenVINO decide which hardware offers the -best performance. For that purpose, just use ``AUTO``. Remember that for -most cases the best hardware is ``GPU`` (better performance, but longer -startup time). +best performance. For that purpose, just use ``AUTO``. + +.. code:: ipython3 + + import ipywidgets as widgets + + core = ov.Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + .. code:: ipython3 # Initialize OpenVINO Runtime. - ie_core = ov.Core() + core = ov.Core() # Read the network and corresponding weights from a file. - model = ie_core.read_model(model=converted_model_path) + model = core.read_model(model=converted_model_path) # Compile the model for CPU (you can choose manually CPU, GPU etc.) # or let the engine choose the best available device (AUTO). - compiled_model = ie_core.compile_model(model=model, device_name="CPU") + compiled_model = core.compile_model(model=model, device_name=device.value) # Get the input and output nodes. input_layer = compiled_model.input(0) @@ -189,11 +254,13 @@ output. -Processing ----------- +Processing `⇑ <#top>`__ +############################################################################################################################### + + +Process Results `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Process Results -~~~~~~~~~~~~~~~ First, list all available classes and create colors for them. Then, in the post-process stage, transform boxes with normalized coordinates @@ -282,8 +349,9 @@ threshold (0.5). Finally, draw boxes and labels inside them. return frame -Main Processing Function -~~~~~~~~~~~~~~~~~~~~~~~~ +Main Processing Function `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Run object detection on the specified source. Either a webcam or a video file. @@ -393,11 +461,13 @@ file. if use_popup: cv2.destroyAllWindows() -Run ---- +Run `⇑ <#top>`__ +############################################################################################################################### + + +Run Live Object Detection `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Run Live Object Detection -~~~~~~~~~~~~~~~~~~~~~~~~~ Use a webcam as the video input. By default, the primary webcam is set with ``source=0``. If you have multiple webcams, each one will be @@ -406,10 +476,12 @@ using a front-facing camera. Some web browsers, especially Mozilla Firefox, may cause flickering. If you experience flickering, set ``use_popup=True``. - **NOTE**: To use this notebook with a webcam, you need to run the - notebook on a computer with a webcam. If you run the notebook on a - server (for example, Binder), the webcam will not work. Popup mode - may not work if you run this notebook on a remote computer (for +.. note:: + + To use this notebook with a webcam, you need to run the + notebook on a computer with a webcam. If you run the notebook on a + server (for example, Binder), the webcam will not work. Popup mode + may not work if you run this notebook on a remote computer (for example, Binder). Run the object detection: @@ -426,12 +498,13 @@ Run the object detection: .. parsed-literal:: - [ WARN:0@43.661] global cap_v4l.cpp:982 open VIDEOIO(V4L2:/dev/video0): can't open camera by index - [ERROR:0@43.661] global obsensor_uvc_stream_channel.cpp:156 getStreamChannelGroup Camera index out of range + [ WARN:0@44.255] global cap_v4l.cpp:982 open VIDEOIO(V4L2:/dev/video0): can't open camera by index + [ERROR:0@44.255] global obsensor_uvc_stream_channel.cpp:156 getStreamChannelGroup Camera index out of range -Run Object Detection on a Video File -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Run Object Detection on a Video File `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + If you do not have a webcam, you can still run this demo with a video file. Any `format supported by @@ -446,7 +519,7 @@ will work. -.. image:: 401-object-detection-with-output_files/401-object-detection-with-output_20_0.png +.. image:: 401-object-detection-with-output_files/401-object-detection-with-output_21_0.png .. parsed-literal:: @@ -454,8 +527,9 @@ will work. Source ended -References ----------- +References `⇑ <#top>`__ +############################################################################################################################### + 1. `SSDLite MobileNetV2 `__ diff --git a/docs/notebooks/401-object-detection-with-output_files/401-object-detection-with-output_20_0.png b/docs/notebooks/401-object-detection-with-output_files/401-object-detection-with-output_20_0.png deleted file mode 100644 index 77d8a125995..00000000000 --- a/docs/notebooks/401-object-detection-with-output_files/401-object-detection-with-output_20_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:0a78530d91ab8d78433364fa1b43b704dbdf50677f535291464302dd32b7d2b7 -size 175073 diff --git a/docs/notebooks/401-object-detection-with-output_files/401-object-detection-with-output_21_0.png b/docs/notebooks/401-object-detection-with-output_files/401-object-detection-with-output_21_0.png new file mode 100644 index 00000000000..8f1c9d1ae95 --- /dev/null +++ b/docs/notebooks/401-object-detection-with-output_files/401-object-detection-with-output_21_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:ae3c173441be8e7cfd682e02f750cb4d02dc0b3678ad37c8c7bb8f41d15d4440 +size 174850 diff --git a/docs/notebooks/401-object-detection-with-output_files/index.html b/docs/notebooks/401-object-detection-with-output_files/index.html index ac21f275ce2..67469b5a0eb 100644 --- a/docs/notebooks/401-object-detection-with-output_files/index.html +++ b/docs/notebooks/401-object-detection-with-output_files/index.html @@ -1,7 +1,7 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/401-object-detection-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/401-object-detection-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/401-object-detection-with-output_files/


../
-401-object-detection-with-output_20_0.png          12-Jul-2023 00:11              175073
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/401-object-detection-with-output_files/


../
+401-object-detection-with-output_21_0.png          16-Aug-2023 01:31              174850
 

diff --git a/docs/notebooks/402-pose-estimation-with-output.rst b/docs/notebooks/402-pose-estimation-with-output.rst index e3383de9cb6..f6f9773a6b3 100644 --- a/docs/notebooks/402-pose-estimation-with-output.rst +++ b/docs/notebooks/402-pose-estimation-with-output.rst @@ -1,6 +1,8 @@ Live Human Pose Estimation with OpenVINO™ ========================================= +.. _top: + This notebook demonstrates live pose estimation with OpenVINO, using the OpenPose `human-pose-estimation-0001 `__ @@ -12,10 +14,31 @@ Additionally, you can also upload a video file. **NOTE**: To use a webcam, you must run this Jupyter notebook on a computer with a webcam. If you run on a server, the webcam will not work. However, you can still do inference on a video in the final - step. + step. + +**Table of contents**: + +- `Imports <#imports>`__ +- `The model <#the-model>`__ + + - `Download the model <#download-the-model>`__ + - `Load the model <#load-the-model>`__ + +- `Processing <#processing>`__ + + - `OpenPose Decoder <#openpose-decoder>`__ + - `Process Results <#process-results>`__ + - `Draw Pose Overlays <#draw-pose-overlays>`__ + - `Main Processing Function <#main-processing-function>`__ + +- `Run <#run>`__ + + - `Run Live Pose Estimation <#run-live-pose-estimation>`__ + - `Run Pose Estimation on a Video File <#run-pose-estimation-on-a-video-file>`__ + +Imports `⇑ <#top>`__ +############################################################################################################################### -Imports -------- .. code:: ipython3 @@ -35,11 +58,13 @@ Imports sys.path.append("../utils") import notebook_utils as utils -The model ---------- +The model `⇑ <#top>`__ +############################################################################################################################### + + +Download the model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Download the model -~~~~~~~~~~~~~~~~~~ Use the ``download_file``, a function from the ``notebook_utils`` file. It automatically creates a directory structure and downloads the @@ -80,8 +105,9 @@ precision in the code below. model/intel/human-pose-estimation-0001/FP16-INT8/human-pose-estimation-0001.bin: 0%| | 0.00/4.03M [… -Load the model -~~~~~~~~~~~~~~ +Load the model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Downloaded models are located in a fixed structure, which indicates a vendor, the name of the model and a precision. @@ -89,16 +115,41 @@ vendor, the name of the model and a precision. Only a few lines of code are required to run the model. First, initialize OpenVINO Runtime. Then, read the network architecture and model weights from the ``.bin`` and ``.xml`` files to compile it for the -desired device. +desired device. Select device from dropdown list for running inference +using OpenVINO. + +.. code:: ipython3 + + import ipywidgets as widgets + + core = Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + .. code:: ipython3 # Initialize OpenVINO Runtime - ie_core = Core() + core = Core() # Read the network from a file. - model = ie_core.read_model(model_path) - # Let the AUTO device decide where to load the model (you can use CPU or GPU). - compiled_model = ie_core.compile_model(model=model, device_name="AUTO", config={"PERFORMANCE_HINT": "LATENCY"}) + model = core.read_model(model_path) + # Let the AUTO device decide where to load the model (you can use CPU, GPU as well). + compiled_model = core.compile_model(model=model, device_name=device.value, config={"PERFORMANCE_HINT": "LATENCY"}) # Get the input and output names of nodes. input_layer = compiled_model.input(0) @@ -109,7 +160,7 @@ desired device. Input layer has the name of the input node and output layers contain names of output nodes of the network. In the case of OpenPose Model, -there is 1 input and 2 outputs: pafs and keypoints heatmap. +there is 1 input and 2 outputs: PAFs and keypoints heatmap. .. code:: ipython3 @@ -124,21 +175,23 @@ there is 1 input and 2 outputs: pafs and keypoints heatmap. -Processing ----------- +Processing `⇑ <#top>`__ +############################################################################################################################### + + +OpenPose Decoder `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -OpenPoseDecoder -~~~~~~~~~~~~~~~ To transform the raw results from the neural network into pose -estimations, you need Open Pose Decoder. It is provided in the `Open +estimations, you need OpenPose Decoder. It is provided in the `Open Model Zoo `__ and compatible with the ``human-pose-estimation-0001`` model. If you choose a model other than ``human-pose-estimation-0001`` you will -need another decoder (for example, AssociativeEmbeddingDecoder), which -is available in the `demos +need another decoder (for example, ``AssociativeEmbeddingDecoder``), +which is available in the `demos section `__ of Open Model Zoo. @@ -146,8 +199,9 @@ of Open Model Zoo. decoder = OpenPoseDecoder() -Process Results -~~~~~~~~~~~~~~~ +Process Results `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + A bunch of useful functions to transform results into poses. @@ -217,8 +271,9 @@ factor. poses[:, :, :2] *= output_scale return poses, scores -Draw Pose Overlays -~~~~~~~~~~~~~~~~~~ +Draw Pose Overlays `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Draw pose overlays on the image to visualize estimated poses. Joints are drawn as circles and limbs are drawn as lines. The code is based on the @@ -255,8 +310,9 @@ from Open Model Zoo. cv2.addWeighted(img, 0.4, img_limbs, 0.6, 0, dst=img) return img -Main Processing Function -~~~~~~~~~~~~~~~~~~~~~~~~ +Main Processing Function `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Run pose estimation on the specified source. Either a webcam or a video file. @@ -350,11 +406,13 @@ file. if use_popup: cv2.destroyAllWindows() -Run ---- +Run `⇑ <#top>`__ +############################################################################################################################### + + +Run Live Pose Estimation `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Run Live Pose Estimation -~~~~~~~~~~~~~~~~~~~~~~~~ Use a webcam as the video input. By default, the primary webcam is set with ``source=0``. If you have multiple webcams, each one will be @@ -383,12 +441,13 @@ Run the pose estimation: .. parsed-literal:: - [ WARN:0@2.610] global cap_v4l.cpp:982 open VIDEOIO(V4L2:/dev/video0): can't open camera by index - [ERROR:0@2.611] global obsensor_uvc_stream_channel.cpp:156 getStreamChannelGroup Camera index out of range + [ WARN:0@2.649] global cap_v4l.cpp:982 open VIDEOIO(V4L2:/dev/video0): can't open camera by index + [ERROR:0@2.649] global obsensor_uvc_stream_channel.cpp:156 getStreamChannelGroup Camera index out of range -Run Pose Estimation on a Video File -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Run Pose Estimation on a Video File `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + If you do not have a webcam, you can still run this demo with a video file. Any `format supported by @@ -403,7 +462,7 @@ will work. You can skip first ``N`` frames to fast forward video. -.. image:: 402-pose-estimation-with-output_files/402-pose-estimation-with-output_20_0.png +.. image:: 402-pose-estimation-with-output_files/402-pose-estimation-with-output_21_0.png .. parsed-literal:: diff --git a/docs/notebooks/402-pose-estimation-with-output_files/402-pose-estimation-with-output_20_0.png b/docs/notebooks/402-pose-estimation-with-output_files/402-pose-estimation-with-output_20_0.png deleted file mode 100644 index 2c108d0a9bd..00000000000 --- a/docs/notebooks/402-pose-estimation-with-output_files/402-pose-estimation-with-output_20_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:82e9d3362fe5e4893dbcd84d8cf13468875b820a8b8683fa11f16b9aaf492b89 -size 107977 diff --git a/docs/notebooks/402-pose-estimation-with-output_files/402-pose-estimation-with-output_21_0.png b/docs/notebooks/402-pose-estimation-with-output_files/402-pose-estimation-with-output_21_0.png new file mode 100644 index 00000000000..450f6ed81d6 --- /dev/null +++ b/docs/notebooks/402-pose-estimation-with-output_files/402-pose-estimation-with-output_21_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:e78567346df5ffd2550196f70571f59da85af3d23cb2d916967b1bcd84e9dd8c +size 107992 diff --git a/docs/notebooks/402-pose-estimation-with-output_files/index.html b/docs/notebooks/402-pose-estimation-with-output_files/index.html index a639781cfb5..a8595cfe92a 100644 --- a/docs/notebooks/402-pose-estimation-with-output_files/index.html +++ b/docs/notebooks/402-pose-estimation-with-output_files/index.html @@ -1,7 +1,7 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/402-pose-estimation-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/402-pose-estimation-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/402-pose-estimation-with-output_files/


../
-402-pose-estimation-with-output_20_0.png           12-Jul-2023 00:11              107977
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/402-pose-estimation-with-output_files/


../
+402-pose-estimation-with-output_21_0.png           16-Aug-2023 01:31              107992
 

diff --git a/docs/notebooks/403-action-recognition-webcam-with-output.rst b/docs/notebooks/403-action-recognition-webcam-with-output.rst index e8b8841c5c1..7a08d4f335d 100644 --- a/docs/notebooks/403-action-recognition-webcam-with-output.rst +++ b/docs/notebooks/403-action-recognition-webcam-with-output.rst @@ -1,6 +1,8 @@ Human Action Recognition with OpenVINO™ ======================================= +.. _top: + This notebook demonstrates live human action recognition with OpenVINO, using the `Action Recognition Models `__ from `Open @@ -18,13 +20,15 @@ notebook shows how to create the following pipeline: Final part of this notebook shows live inference results from a webcam. Additionally, you can also upload a video file. -**NOTE**: To use a webcam, you must run this Jupyter notebook on a -computer with a webcam. If you run on a server, the webcam will not -work. However, you can still do inference on a video in the final step. +.. note:: + + To use a webcam, you must run this Jupyter notebook on a computer with a webcam. + If you run on a server, the webcam will not work. However, you can still do + inference on a video in the final step. -------------- -[1] seq2seq: Deep learning models that take a sequence of items to the +[1] ``seq2seq``: Deep learning models that take a sequence of items to the input and output. In this case, input: video frames, output: actions sequence. This ``"seq2seq"`` is composed of an encoder and a decoder. The encoder captures ``"context"`` of the inputs to be analyzed by the @@ -35,8 +39,27 @@ Transformer and `ResNet34 `__. -Imports -------- +**Table of contents**: + +- `Imports <#imports>`__ +- `The models <#the-models>`__ + + - `Download the models <#download-the-models>`__ + - `Load your labels <#load-your-labels>`__ + - `Load the models <#load-the-models>`__ + + - `Model Initialization function <#model-initialization-function>`__ + - `Initialization for Encoder and Decoder <#initialization-for-encoder-and-decoder>`__ + + - `Helper functions <#helper-functions>`__ + - `AI Functions <#ai-functions>`__ + - `Main Processing Function <#main-processing-function>`__ + - `Run Action Recognition on a Video File <#run-action-recognition-on-a-video-file>`__ + - `Run Action Recognition Using a Webcam <#run-action-recognition-using-a-webcam>`__ + +Imports `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -55,11 +78,13 @@ Imports sys.path.append("../utils") import notebook_utils as utils -The models ----------- +The models `⇑ <#top>`__ +############################################################################################################################### + + +Download the models `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Download the models -~~~~~~~~~~~~~~~~~~~ Use ``omz_downloader``, which is a command-line tool from the ``openvino-dev`` package. It automatically creates a directory structure @@ -70,10 +95,11 @@ and the system automatically downloads the two models ``"action-recognition-0001-encoder"`` and ``"action-recognition-0001-decoder"`` - **NOTE**: If you want to download another model, such as - ``"driver-action-recognition-adas-0002"`` - (``"driver-action-recognition-adas-0002-encoder"`` + - ``"driver-action-recognition-adas-0002-decoder"``), replace the name +.. note:: + + If you want to download another model, such as + ``"driver-action-recognition-adas-0002"`` (``"driver-action-recognition-adas-0002-encoder"`` + + ``"driver-action-recognition-adas-0002-decoder"``), replace the name of the model in the code below. Using a model outside the list can require different pre- and post-processing. @@ -119,8 +145,9 @@ and the system automatically downloads the two models -Load your labels -~~~~~~~~~~~~~~~~ +Load your labels `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + This tutorial uses `Kinetics-400 dataset `__, and @@ -145,8 +172,9 @@ also provides the text file embedded into this notebook. ['abseiling', 'air drumming', 'answering questions', 'applauding', 'applying cream', 'archery', 'arm wrestling', 'arranging flowers', 'assembling computer'] (400,) -Load the models -~~~~~~~~~~~~~~~ +Load the models `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Load the two models for this particular architecture, Encoder and Decoder. Downloaded models are located in a fixed structure, indicating @@ -155,26 +183,54 @@ a vendor, the name of the model, and a precision. 1. Initialize OpenVINO Runtime. 2. Read the network from ``*.bin`` and ``*.xml`` files (weights and architecture). -3. Compile the model for CPU. +3. Compile the model for specified device. 4. Get input and output names of nodes. Only a few lines of code are required to run the model. -Model Initialization function -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +Select device from dropdown list for running inference using OpenVINO + +.. code:: ipython3 + + import ipywidgets as widgets + + core = Core() + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +Model Initialization function `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + .. code:: ipython3 # Initialize OpenVINO Runtime. - ie_core = Core() + core = Core() - def model_init(model_path: str) -> Tuple: + def model_init(model_path: str, device: str) -> Tuple: """ Read the network and weights from a file, load the model on CPU and get input and output names of nodes - :param: model: model architecture path *.xml + :param: + model: model architecture path *.xml + device: inference device :retuns: compiled_model: Compiled model input_key: Input node for model @@ -182,31 +238,33 @@ Model Initialization function """ # Read the network and corresponding weights from a file. - model = ie_core.read_model(model=model_path) - # Compile the model for CPU (you can also use GPU). - compiled_model = ie_core.compile_model(model=model, device_name="CPU") + model = core.read_model(model=model_path) + # Compile the model for specified device. + compiled_model = core.compile_model(model=model, device_name=device) # Get input and output names of nodes. input_keys = compiled_model.input(0) output_keys = compiled_model.output(0) return input_keys, output_keys, compiled_model -Initialization for Encoder and Decoder -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +Initialization for Encoder and Decoder `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + .. code:: ipython3 # Encoder initialization - input_key_en, output_keys_en, compiled_model_en = model_init(model_path_encoder) + input_key_en, output_keys_en, compiled_model_en = model_init(model_path_encoder, device.value) # Decoder initialization - input_key_de, output_keys_de, compiled_model_de = model_init(model_path_decoder) + input_key_de, output_keys_de, compiled_model_de = model_init(model_path_decoder, device.value) # Get input size - Encoder. height_en, width_en = list(input_key_en.shape)[2:] # Get input size - Decoder. frames2decode = list(input_key_de.shape)[0:][1] -Helper functions -~~~~~~~~~~~~~~~~ +Helper functions `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Use the following helper functions for preprocessing and postprocessing frames: @@ -326,8 +384,9 @@ frames: cv2.putText(frame, display_text, text_loc2, FONT_STYLE, FONT_SIZE, FONT_COLOR2) cv2.putText(frame, display_text, text_loc, FONT_STYLE, FONT_SIZE, FONT_COLOR) -AI Functions -~~~~~~~~~~~~ +AI Functions `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Following the pipeline above, you will use the next functions to: @@ -416,8 +475,9 @@ Following the pipeline above, you will use the next functions to: exp = np.exp(x) return exp / np.sum(exp, axis=None) -Main Processing Function -~~~~~~~~~~~~~~~~~~~~~~~~ +Main Processing Function `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Running action recognition function will run in different operations, either a webcam or a video file. See the list of procedures below: @@ -567,8 +627,9 @@ either a webcam or a video file. See the list of procedures below: if use_popup: cv2.destroyAllWindows() -Run Action Recognition on a Video File -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Run Action Recognition on a Video File `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Find out how the model works in a video file. `Any format supported `__ @@ -576,10 +637,11 @@ by OpenCV will work. You can press the stop button anytime while the video file is running, and it will activate the webcam for the next step. - **NOTE**: Sometimes, the video can be cut off if there are corrupted - frames. In that case, you can convert it. If you experience any - problems with your video, use the - `HandBrake `__ and select the MPEG format. +.. note:: + + Sometimes, the video can be cut off if there are corrupted frames. In that + case, you can convert it. If you experience any problems with your video, + use the `HandBrake `__ and select the MPEG format. .. code:: ipython3 @@ -588,7 +650,7 @@ step. -.. image:: 403-action-recognition-webcam-with-output_files/403-action-recognition-webcam-with-output_19_0.png +.. image:: 403-action-recognition-webcam-with-output_files/403-action-recognition-webcam-with-output_21_0.png .. parsed-literal:: @@ -596,8 +658,9 @@ step. Source ended -Run Action Recognition Using a Webcam -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Run Action Recognition Using a Webcam `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Now, try to see yourself in your webcam. @@ -618,6 +681,6 @@ Now, try to see yourself in your webcam. .. parsed-literal:: - [ WARN:0@318.296] global cap_v4l.cpp:982 open VIDEOIO(V4L2:/dev/video0): can't open camera by index - [ERROR:0@318.297] global obsensor_uvc_stream_channel.cpp:156 getStreamChannelGroup Camera index out of range + [ WARN:0@319.035] global cap_v4l.cpp:982 open VIDEOIO(V4L2:/dev/video0): can't open camera by index + [ERROR:0@319.035] global obsensor_uvc_stream_channel.cpp:156 getStreamChannelGroup Camera index out of range diff --git a/docs/notebooks/403-action-recognition-webcam-with-output_files/403-action-recognition-webcam-with-output_19_0.png b/docs/notebooks/403-action-recognition-webcam-with-output_files/403-action-recognition-webcam-with-output_19_0.png deleted file mode 100644 index 90314429a8e..00000000000 --- a/docs/notebooks/403-action-recognition-webcam-with-output_files/403-action-recognition-webcam-with-output_19_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:f3b5d63d24428121a508462045e8fcb42970bb1e94b14a86e62092f7f6fc8364 -size 67618 diff --git a/docs/notebooks/403-action-recognition-webcam-with-output_files/403-action-recognition-webcam-with-output_21_0.png b/docs/notebooks/403-action-recognition-webcam-with-output_files/403-action-recognition-webcam-with-output_21_0.png new file mode 100644 index 00000000000..c1b7138364d --- /dev/null +++ b/docs/notebooks/403-action-recognition-webcam-with-output_files/403-action-recognition-webcam-with-output_21_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:789f37b6428ab4468a32c7ea70838dcc6eeea93650647a75d4202d814e5af43e +size 68035 diff --git a/docs/notebooks/403-action-recognition-webcam-with-output_files/index.html b/docs/notebooks/403-action-recognition-webcam-with-output_files/index.html index ca196d45e04..80637667b61 100644 --- a/docs/notebooks/403-action-recognition-webcam-with-output_files/index.html +++ b/docs/notebooks/403-action-recognition-webcam-with-output_files/index.html @@ -1,7 +1,7 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/403-action-recognition-webcam-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/403-action-recognition-webcam-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/403-action-recognition-webcam-with-output_files/


../
-403-action-recognition-webcam-with-output_19_0.png 12-Jul-2023 00:11               67618
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/403-action-recognition-webcam-with-output_files/


../
+403-action-recognition-webcam-with-output_21_0.png 16-Aug-2023 01:31               68035
 

diff --git a/docs/notebooks/404-style-transfer-with-output.rst b/docs/notebooks/404-style-transfer-with-output.rst index f9f24064362..4854386268c 100644 --- a/docs/notebooks/404-style-transfer-with-output.rst +++ b/docs/notebooks/404-style-transfer-with-output.rst @@ -1,6 +1,8 @@ Style Transfer with OpenVINO™ ============================= +.. _top: + This notebook demonstrates style transfer with OpenVINO, using the Style Transfer Models from `ONNX Model Repository `__. Specifically, `Fast @@ -23,16 +25,40 @@ and Super-Resolution `__ along with part of this notebook shows live inference results from a webcam. Additionally, you can also upload a video file. - **NOTE**: If you have a webcam on your computer, you can see live - results streaming in the notebook. If you run the notebook on a - server, the webcam will not work but you can run inference, using a - video file. +.. note:: -Preparation ------------ + If you have a webcam on your computer, you can see live results streaming in + the notebook. If you run the notebook on a server, the webcam will not work + but you can run inference, using a video file. + + +**Table of contents**: + +- `Preparation <#preparation>`__ + + - `Install requirements <#install-requirements>`__ + - `Imports <#imports>`__ + +- `The Model <#the-model>`__ + + - `Download the Model <#download-the-model>`__ + - `Convert ONNX Model to OpenVINO IR Format <#convert-onnx-model-to-openvino-ir-format>`__ + - `Load the Model <#load-the-model>`__ + - `Preprocess the image <#preprocess-the-image>`__ + - `Helper function to postprocess the stylized image <#helper-function-to-postprocess-the-stylized-image>`__ + - `Main Processing Function <#main-processing-function>`__ + - `Run Style Transfer Using a Webcam <#run-style-transfer-using-a-webcam>`__ + - `Run Style Transfer on a Video File <#run-style-transfer-on-a-video-file>`__ + +- `References <#references>`__ + +Preparation `⇑ <#top>`__ +############################################################################################################################### + + +Install requirements `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Install requirements -~~~~~~~~~~~~~~~~~~~~ .. code:: ipython3 @@ -46,8 +72,9 @@ Install requirements filename='notebook_utils.py' ) -Imports -~~~~~~~ +Imports `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -77,11 +104,13 @@ Pointilism to do the style transfer. interactive(lambda option: print(option), option=styleButtons) -The Model ---------- +The Model `⇑ <#top>`__ +############################################################################################################################### + + +Download the Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Download the Model -~~~~~~~~~~~~~~~~~~ The style transfer model, selected in the previous step, will be downloaded to ``model_path`` if you have not already downloaded it. The @@ -102,23 +131,24 @@ OpenVINO Intermediate Representation (IR) with ``FP16`` precision. style_url = f"{base_url}/{model_path}" utils.download_file(style_url, directory=base_model_dir) -Convert ONNX Model to OpenVINO IR Format -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Convert ONNX Model to OpenVINO IR Format `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + In the next step, you will convert the ONNX model to OpenVINO IR format with ``FP16`` precision. While ONNX models are directly supported by OpenVINO runtime, it can be useful to convert them to IR format to take advantage of OpenVINO optimization tools and features. The -``mo.convert_model`` python function can be used for converting model, -using OpenVINO Model Optimizer. The converted model is saved to the -model directory. The function returns instance of OpenVINO Model class, -which is ready to use in Python interface but can also be serialized to -OpenVINO IR format for future execution. If the model has been already -converted, you can skip this step. +``mo.convert_model`` Python function of model conversion API can be +used. The converted model is saved to the model directory. The function +returns instance of OpenVINO Model class, which is ready to use in +Python interface but can also be serialized to OpenVINO IR format for +future execution. If the model has been already converted, you can skip +this step. .. code:: ipython3 - # Construct the command for Model Optimizer. + # Construct the command for model conversion API. from openvino.runtime import serialize from openvino.tools import mo @@ -131,8 +161,9 @@ converted, you can skip this step. ir_path = Path(f"model/{styleButtons.value.lower()}-9.xml") onnx_path = Path(f"model/{model_path}") -Load the Model -~~~~~~~~~~~~~~ +Load the Model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Both the ONNX model(s) and converted IR model(s) are stored in the ``model`` directory. @@ -145,7 +176,8 @@ it to load, as the startup time is somewhat longer than ``CPU``. To let OpenVINO automatically select the best device for inference just use ``AUTO``. In most cases, the best device to use is ``GPU`` (better -performance, but slightly longer startup time). +performance, but slightly longer startup time). You can select one from +available devices using dropdown list below. OpenVINO Runtime can load ONNX models from `ONNX Model Repository `__ directly. In such cases, @@ -156,17 +188,33 @@ results. .. code:: ipython3 # Initialize OpenVINO Runtime. - ie_core = Core() + core = Core() # Read the network and corresponding weights from ONNX Model. # model = ie_core.read_model(model=onnx_path) # Read the network and corresponding weights from IR Model. - model = ie_core.read_model(model=ir_path) + model = core.read_model(model=ir_path) + +.. code:: ipython3 + + import ipywidgets as widgets + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + # Compile the model for CPU (or change to GPU, etc. for other devices) # or let OpenVINO select the best available device with AUTO. - compiled_model = ie_core.compile_model(model=model, device_name="AUTO") + device + +.. code:: ipython3 + + compiled_model = core.compile_model(model=model, device_name=device.value) # Get the input and output nodes. input_layer = compiled_model.input(0) @@ -185,12 +233,11 @@ respectively. For *fast-neural-style-mosaic-onnx*, there is 1 input and # Get the input size. N, C, H, W = list(input_layer.shape) -Preprocess the image -~~~~~~~~~~~~~~~~~~~~ +Preprocess the image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Preprocess the input image before running the model. Prepare the -dimensions and channel order for the image to match the original image -with the input tensor +Preprocess the input image before running the model. Prepare the dimensions and channel order for the +image to match the original image with the input tensor 1. Preprocess a frame to convert from ``unit8`` to ``float32``. 2. Transpose the array to match with the network input size @@ -215,12 +262,11 @@ with the input tensor image = np.expand_dims(image, axis=0) return image -Helper function to postprocess the stylized image -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Helper function to postprocess the stylized image `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -The converted IR model outputs a NumPy ``float32`` array of the `(1, 3, -224, -224) `__ +The converted IR model outputs a NumPy ``float32`` array of the +`(1, 3, 224,224) `__ shape . .. code:: ipython3 @@ -242,8 +288,9 @@ shape . stylized_image = cv2.cvtColor(stylized_image, cv2.COLOR_BGR2RGB) return stylized_image -Main Processing Function -~~~~~~~~~~~~~~~~~~~~~~~~ +Main Processing Function `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + The style transfer function can be run in different operating modes, either using a webcam or a video file. @@ -339,8 +386,9 @@ either using a webcam or a video file. if use_popup: cv2.destroyAllWindows() -Run Style Transfer Using a Webcam -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Run Style Transfer Using a Webcam `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Now, try to apply the style transfer model using video from your webcam. By default, the primary webcam is set with ``source=0``. If you have @@ -358,19 +406,22 @@ experience flickering, set ``use_popup=True``. run_style_transfer(source=0, flip=True, use_popup=False) -Run Style Transfer on a Video File -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Run Style Transfer on a Video File `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + You can find out how the model works with a video file. For that, use -any `formats supported by -OpenCV `__. +any `formats supported by OpenCV `__. You can press the stop button to terminate anytime while the video file is running. - **NOTE**: Sometimes, the video will be cut off when frames are - corrupted. If this happens, or you experience any other problems with - your video, use the `HandBrake `__ encoder - tool to create a video file in MPEG format. +.. note:: + + Sometimes, the video will be cut off when frames are corrupted. If this + happens, or you experience any other problems with your video, use the + `HandBrake `__ encoder tool to create a video file in + MPEG format. + .. code:: ipython3 @@ -379,7 +430,7 @@ is running. -.. image:: 404-style-transfer-with-output_files/404-style-transfer-with-output_25_0.png +.. image:: 404-style-transfer-with-output_files/404-style-transfer-with-output_27_0.png .. parsed-literal:: @@ -387,8 +438,9 @@ is running. Source ended -References ----------- +References `⇑ <#top>`__ +############################################################################################################################### + 1. `ONNX Model Zoo `__ 2. `Fast Neural Style diff --git a/docs/notebooks/404-style-transfer-with-output_files/404-style-transfer-with-output_25_0.png b/docs/notebooks/404-style-transfer-with-output_files/404-style-transfer-with-output_27_0.png similarity index 100% rename from docs/notebooks/404-style-transfer-with-output_files/404-style-transfer-with-output_25_0.png rename to docs/notebooks/404-style-transfer-with-output_files/404-style-transfer-with-output_27_0.png diff --git a/docs/notebooks/404-style-transfer-with-output_files/index.html b/docs/notebooks/404-style-transfer-with-output_files/index.html index 67657087cf8..a8ea8167a1f 100644 --- a/docs/notebooks/404-style-transfer-with-output_files/index.html +++ b/docs/notebooks/404-style-transfer-with-output_files/index.html @@ -1,7 +1,7 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/404-style-transfer-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/404-style-transfer-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/404-style-transfer-with-output_files/


../
-404-style-transfer-with-output_25_0.png            12-Jul-2023 00:11               81832
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/404-style-transfer-with-output_files/


../
+404-style-transfer-with-output_27_0.png            16-Aug-2023 01:31               81832
 

diff --git a/docs/notebooks/405-paddle-ocr-webcam-with-output.rst b/docs/notebooks/405-paddle-ocr-webcam-with-output.rst index fb7bb2cc9de..608a9d4ab58 100644 --- a/docs/notebooks/405-paddle-ocr-webcam-with-output.rst +++ b/docs/notebooks/405-paddle-ocr-webcam-with-output.rst @@ -1,11 +1,13 @@ PaddleOCR with OpenVINO™ ======================== +.. _top: + This demo shows how to run PP-OCR model on OpenVINO natively. Instead of exporting the PaddlePaddle model to ONNX and then converting to the -OpenVINO Intermediate Representation (OpenVINO IR) format with Model -Optimizer, you can now read directly from the PaddlePaddle Model without -any conversions. +OpenVINO Intermediate Representation (OpenVINO IR) format with model +conversion API, you can now read directly from the PaddlePaddle Model +without any conversions. `PaddleOCR `__ is an ultra-light OCR model trained with PaddlePaddle deep learning framework, that aims to create multilingual and practical OCR tools. @@ -13,23 +15,51 @@ that aims to create multilingual and practical OCR tools. The PaddleOCR pre-trained model used in the demo refers to the *“Chinese and English ultra-lightweight PP-OCR model (9.4M)”*. More open source pre-trained models can be downloaded at `PaddleOCR -Github `__ or `PaddleOCR +GitHub `__ or `PaddleOCR Gitee `__. Working pipeline of the PaddleOCR is as follows: - **NOTE**: To use this notebook with a webcam, you need to run the - notebook on a computer with a webcam. If you run the notebook on a - server, the webcam will not work. You can still do inference on a - video file. +.. note:: + + To use this notebook with a webcam, you need to run the notebook on a computer + with a webcam. If you run the notebook on a server, the webcam will not work. + You can still do inference on a video file. + +**Table of contents**: + +- `Imports <#imports>`__ + + - `Select inference device <#select-inference-device>`__ + - `Models for PaddleOCR <#models-for-paddleocr>`__ + + - `Download the Model for Text Detection <#download-the-model-for-text-detection>`__ + - `Load the Model for Text Detection <#load-the-model-for-text-detection>`__ + - `Download the Model for Text Recognition <#download-the-model-for-text-recognition>`__ + - `Load the Model for Text Recognition with Dynamic Shape <#load-the-model-for-text-recognition-with-dynamic-shape>`__ + + - `Preprocessing Image Functions for Text Detection and Recognition <#preprocessing-image-functions-for-text-detection-and-recognition>`__ + - `Postprocessing Image for Text Detection <#postprocessing-image-for-text-detection>`__ + - `Main Processing Function for PaddleOCR <#main-processing-function-for-paddleocr>`__ + +- `Run Live PaddleOCR with OpenVINO <#run-live-paddleocr-with-openvino>`__ .. code:: ipython3 - !pip install -q 'openvino-dev>=2023.0.0' - !pip install -q 'paddlepaddle==2.5.0rc0' - !pip install -q 'pyclipper>=1.2.1' 'shapely>=1.7.1' + !pip install -q "openvino-dev>=2023.0.0" + !pip install -q "paddlepaddle==2.5.0" + !pip install -q "pyclipper>=1.2.1" "shapely>=1.7.1" + + +.. parsed-literal:: + + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + + +Imports `⇑ <#top>`__ +############################################################################################################################### -Imports -------- .. code:: ipython3 @@ -66,8 +96,39 @@ Imports import notebook_utils as utils import pre_post_processing as processing -Models for PaddleOCR -~~~~~~~~~~~~~~~~~~~~ +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + + core = Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +Models for PaddleOCR `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + PaddleOCR includes two parts of deep learning models, text detection and text recognition. Pre-trained models used in the demo are downloaded and @@ -110,14 +171,15 @@ files to load to CPU/GPU. else: print("Error Extracting the model. Please check the network.") -Download the Model for Text **Detection** -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +Download the Model for Text **Detection** `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + .. code:: ipython3 # A directory where the model will be downloaded. - det_model_url = "https://paddleocr.bj.bcebos.com/PP-OCRv3/chinese/ch_PP-OCRv3_det_infer.tar" + det_model_url = "https://storage.openvinotoolkit.org/repositories/openvino_notebooks/models/paddle-ocr/ch_PP-OCRv3_det_infer.tar" det_model_file_path = Path("model/ch_PP-OCRv3_det_infer/inference.pdmodel") run_model_download(det_model_url, det_model_file_path) @@ -131,7 +193,7 @@ Download the Model for Text **Detection** .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/405-padd… + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/405-padd… .. parsed-literal:: @@ -140,26 +202,28 @@ Download the Model for Text **Detection** Model Extracted to model/ch_PP-OCRv3_det_infer/inference.pdmodel. -Load the Model for Text **Detection** -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +Load the Model for Text **Detection** `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + .. code:: ipython3 # Initialize OpenVINO Runtime for text detection. core = Core() det_model = core.read_model(model=det_model_file_path) - det_compiled_model = core.compile_model(model=det_model, device_name="AUTO") + det_compiled_model = core.compile_model(model=det_model, device_name=device.value) # Get input and output nodes for text detection. det_input_layer = det_compiled_model.input(0) det_output_layer = det_compiled_model.output(0) -Download the Model for Text **Recognition** -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +Download the Model for Text **Recognition** `⇑ <#top>`__ +------------------------------------------------------------------------------------------------------------------------------- + .. code:: ipython3 - rec_model_url = "https://paddleocr.bj.bcebos.com/PP-OCRv3/chinese/ch_PP-OCRv3_rec_infer.tar" + rec_model_url = "https://storage.openvinotoolkit.org/repositories/openvino_notebooks/models/paddle-ocr/ch_PP-OCRv3_rec_infer.tar" rec_model_file_path = Path("model/ch_PP-OCRv3_rec_infer/inference.pdmodel") run_model_download(rec_model_url, rec_model_file_path) @@ -173,7 +237,7 @@ Download the Model for Text **Recognition** .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/405-padd… + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/405-padd… .. parsed-literal:: @@ -183,7 +247,7 @@ Download the Model for Text **Recognition** Load the Model for Text **Recognition** with Dynamic Shape -^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +`⇑ <#top>`__ Input to text recognition model refers to detected bounding boxes with different image sizes, for example, dynamic input shapes. Hence: @@ -211,10 +275,11 @@ different image sizes, for example, dynamic input shapes. Hence: rec_input_layer = rec_compiled_model.input(0) rec_output_layer = rec_compiled_model.output(0) -Preprocessing Image Functions for Text Detection and Recognition -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Preprocessing Image Functions for Text Detection and Recognition. `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Define preprosessing functions for text detection and recognition: 1. + +Define preprocessing functions for text detection and recognition: 1. Preprocessing for text detection: resize and normalize input images. 2. Preprocessing for text recognition: resize and normalize detected box images to the same size (for example, ``(3, 32, 320)`` size for images @@ -327,8 +392,9 @@ with Chinese text) for easy batching in inference. norm_img_batch = norm_img_batch.copy() return norm_img_batch -Postprocessing Image for Text Detection -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Postprocessing Image for Text Detection `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -366,8 +432,9 @@ Postprocessing Image for Text Detection dt_boxes = processing.filter_tag_det_res(dt_boxes, ori_im.shape) return dt_boxes -Main Processing Function for PaddleOCR -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Main Processing Function for PaddleOCR `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Run ``paddleOCR`` function in different operations, either a webcam or a video file. See the list of procedures below: @@ -539,8 +606,9 @@ video file. See the list of procedures below: if use_popup: cv2.destroyAllWindows() -Run Live PaddleOCR with OpenVINO --------------------------------- +Run Live PaddleOCR with OpenVINO `⇑ <#top>`__ +############################################################################################################################### + Use a webcam as the video input. By default, the primary webcam is set with ``source=0``. If you have multiple webcams, each one will be @@ -549,8 +617,10 @@ using a front-facing camera. Some web browsers, especially Mozilla Firefox, may cause flickering. If you experience flickering, set ``use_popup=True``. - **NOTE**: Popup mode may not work if you run this notebook on a - remote computer. +.. note:: + + Popup mode may not work if you run this notebook on a remote computer. + Run live PaddleOCR: @@ -566,13 +636,12 @@ Run live PaddleOCR: .. parsed-literal:: - [ WARN:0@49.013] global cap_v4l.cpp:982 open VIDEOIO(V4L2:/dev/video0): can't open camera by index - [ERROR:0@49.013] global obsensor_uvc_stream_channel.cpp:156 getStreamChannelGroup Camera index out of range + [ WARN:0@10.144] global cap_v4l.cpp:982 open VIDEOIO(V4L2:/dev/video0): can't open camera by index + [ERROR:0@10.145] global obsensor_uvc_stream_channel.cpp:156 getStreamChannelGroup Camera index out of range If you do not have a webcam, you can still run this demo with a video -file. Any `format supported by -OpenCV `__ +file. Any `format supported by OpenCV `__ will work. .. code:: ipython3 @@ -584,7 +653,7 @@ will work. -.. image:: 405-paddle-ocr-webcam-with-output_files/405-paddle-ocr-webcam-with-output_30_0.png +.. image:: 405-paddle-ocr-webcam-with-output_files/405-paddle-ocr-webcam-with-output_32_0.png .. parsed-literal:: diff --git a/docs/notebooks/405-paddle-ocr-webcam-with-output_files/405-paddle-ocr-webcam-with-output_30_0.png b/docs/notebooks/405-paddle-ocr-webcam-with-output_files/405-paddle-ocr-webcam-with-output_30_0.png deleted file mode 100644 index c9264893506..00000000000 --- a/docs/notebooks/405-paddle-ocr-webcam-with-output_files/405-paddle-ocr-webcam-with-output_30_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:35014dbf6e20e30909ea8987d6cb466a28de997ff0ba541b05395e3eb3442a1a -size 586573 diff --git a/docs/notebooks/405-paddle-ocr-webcam-with-output_files/405-paddle-ocr-webcam-with-output_32_0.png b/docs/notebooks/405-paddle-ocr-webcam-with-output_files/405-paddle-ocr-webcam-with-output_32_0.png new file mode 100644 index 00000000000..972ebfac5c7 --- /dev/null +++ b/docs/notebooks/405-paddle-ocr-webcam-with-output_files/405-paddle-ocr-webcam-with-output_32_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:a31272403eaeed49ca89304491a237c433dc80dd497c1b813a385baa0056e09f +size 590865 diff --git a/docs/notebooks/405-paddle-ocr-webcam-with-output_files/index.html b/docs/notebooks/405-paddle-ocr-webcam-with-output_files/index.html index efb77092146..103c98d3c8c 100644 --- a/docs/notebooks/405-paddle-ocr-webcam-with-output_files/index.html +++ b/docs/notebooks/405-paddle-ocr-webcam-with-output_files/index.html @@ -1,7 +1,7 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/405-paddle-ocr-webcam-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/405-paddle-ocr-webcam-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/405-paddle-ocr-webcam-with-output_files/


../
-405-paddle-ocr-webcam-with-output_30_0.png         12-Jul-2023 00:11              586573
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/405-paddle-ocr-webcam-with-output_files/


../
+405-paddle-ocr-webcam-with-output_32_0.png         16-Aug-2023 01:31              590865
 

diff --git a/docs/notebooks/406-3D-pose-estimation-with-output.rst b/docs/notebooks/406-3D-pose-estimation-with-output.rst index 7160bf1a6cc..82a00c18827 100644 --- a/docs/notebooks/406-3D-pose-estimation-with-output.rst +++ b/docs/notebooks/406-3D-pose-estimation-with-output.rst @@ -1,6 +1,8 @@ Live 3D Human Pose Estimation with OpenVINO =========================================== +.. _top: + This notebook demonstrates live 3D Human Pose Estimation with OpenVINO via a webcam. We utilize the model `human-pose-estimation-3d-0001 `__ @@ -11,25 +13,52 @@ of this notebook, you will see live inference results from your webcam out the algorithms. **Make sure you have properly installed the**\ `Jupyter extension `__\ **and -been using Jupyterlab to run the demo as suggested in the README.md** +been using JupyterLab to run the demo as suggested in the +``README.md``** **NOTE**: *To use a webcam, you must run this Jupyter notebook on a computer with a webcam. If you run on a remote server, the webcam will not work. However, you can still do inference on a video file in - the final step. This demo utilizes the Python interface in Three.js - integrated with WebGL to process data from the model inference. These - results are processed and displayed in the notebook.* + the final step. This demo utilizes the Python interface in + ``Three.js`` integrated with WebGL to process data from the model + inference. These results are processed and displayed in the + notebook.* *To ensure that the results are displayed correctly, run the code in a recommended browser on one of the following operating systems:* *Ubuntu, Windows: Chrome* *macOS: Safari* -Prerequisites -------------- +**Table of contents**: -**The Pythreejs extension may not display properly when using the latest -Jupyter Notebook release (2.4.1). Therefore, it is recommended to use -Jupyter Lab instead.** +- `Prerequisites <#prerequisites>`__ +- `Imports <#imports>`__ +- `The model <#the-model>`__ + + - `Download the model <#download-the-model>`__ + - `Convert Model to OpenVINO IR format <#convert-model-to-openvino-ir-format>`__ + - `Select inference device <#select-inference-device>`__ + - `Load the model <#load-the-model>`__ + +- `Processing <#processing>`__ + + - `Model Inference <#model-inference>`__ + - `Draw 2D Pose Overlays <#draw-2d-pose-overlays>`__ + - `Main Processing Function <#main-processing-function>`__ + +- `Run <#run>`__ + + - `Run Live Pose Estimation <#run-live-pose-estimation>`__ + - `Run Pose Estimation on a Video File <#run-pose-estimation-on-a-video-file>`__ + +Prerequisites `⇑ <#top>`__ +############################################################################################################################### + + +.. note:: + + The ``pythreejs`` extension may not display properly when using the latest + Jupyter Notebook release (2.4.1). Therefore, it is recommended to use + Jupyter Lab instead. .. code:: ipython3 @@ -40,53 +69,44 @@ Jupyter Lab instead.** Collecting pythreejs Using cached pythreejs-2.4.2-py3-none-any.whl (3.4 MB) - Requirement already satisfied: ipywidgets>=7.2.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from pythreejs) (8.0.7) + Requirement already satisfied: ipywidgets>=7.2.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from pythreejs) (8.1.0) Collecting ipydatawidgets>=1.1.1 (from pythreejs) - Using cached ipydatawidgets-4.3.5-py2.py3-none-any.whl (271 kB) - Requirement already satisfied: numpy in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from pythreejs) (1.23.5) - Requirement already satisfied: traitlets in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from pythreejs) (5.9.0) + Obtaining dependency information for ipydatawidgets>=1.1.1 from https://files.pythonhosted.org/packages/f1/5b/e63c877c4c94382b66de5045e08ec8cd960e8a4d22f0d62a4dfb1f9e5ac6/ipydatawidgets-4.3.5-py2.py3-none-any.whl.metadata + Using cached ipydatawidgets-4.3.5-py2.py3-none-any.whl.metadata (1.4 kB) + Requirement already satisfied: numpy in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from pythreejs) (1.23.5) + Requirement already satisfied: traitlets in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from pythreejs) (5.9.0) Collecting traittypes>=0.2.0 (from ipydatawidgets>=1.1.1->pythreejs) Using cached traittypes-0.2.1-py2.py3-none-any.whl (8.6 kB) - Requirement already satisfied: ipykernel>=4.5.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipywidgets>=7.2.1->pythreejs) (6.24.0) - Requirement already satisfied: ipython>=6.1.0 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipywidgets>=7.2.1->pythreejs) (8.12.2) - Requirement already satisfied: widgetsnbextension~=4.0.7 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipywidgets>=7.2.1->pythreejs) (4.0.8) - Requirement already satisfied: jupyterlab-widgets~=3.0.7 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipywidgets>=7.2.1->pythreejs) (3.0.8) - Requirement already satisfied: comm>=0.1.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipykernel>=4.5.1->ipywidgets>=7.2.1->pythreejs) (0.1.3) - Requirement already satisfied: debugpy>=1.6.5 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipykernel>=4.5.1->ipywidgets>=7.2.1->pythreejs) (1.6.7) - Requirement already satisfied: jupyter-client>=6.1.12 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipykernel>=4.5.1->ipywidgets>=7.2.1->pythreejs) (8.3.0) - Requirement already satisfied: jupyter-core!=5.0.*,>=4.12 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipykernel>=4.5.1->ipywidgets>=7.2.1->pythreejs) (5.3.1) - Requirement already satisfied: matplotlib-inline>=0.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipykernel>=4.5.1->ipywidgets>=7.2.1->pythreejs) (0.1.6) - Requirement already satisfied: nest-asyncio in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipykernel>=4.5.1->ipywidgets>=7.2.1->pythreejs) (1.5.6) - Requirement already satisfied: packaging in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipykernel>=4.5.1->ipywidgets>=7.2.1->pythreejs) (23.1) - Requirement already satisfied: psutil in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipykernel>=4.5.1->ipywidgets>=7.2.1->pythreejs) (5.9.5) - Requirement already satisfied: pyzmq>=20 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipykernel>=4.5.1->ipywidgets>=7.2.1->pythreejs) (25.1.0) - Requirement already satisfied: tornado>=6.1 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipykernel>=4.5.1->ipywidgets>=7.2.1->pythreejs) (6.3.2) - Requirement already satisfied: backcall in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (0.2.0) - Requirement already satisfied: decorator in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (4.4.2) - Requirement already satisfied: jedi>=0.16 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (0.18.2) - Requirement already satisfied: pickleshare in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (0.7.5) - Requirement already satisfied: prompt-toolkit!=3.0.37,<3.1.0,>=3.0.30 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (3.0.39) - Requirement already satisfied: pygments>=2.4.0 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (2.15.1) - Requirement already satisfied: stack-data in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (0.6.2) - Requirement already satisfied: typing-extensions in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (4.7.1) - Requirement already satisfied: pexpect>4.3 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (4.8.0) - Requirement already satisfied: parso<0.9.0,>=0.8.0 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from jedi>=0.16->ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (0.8.3) - Requirement already satisfied: importlib-metadata>=4.8.3 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from jupyter-client>=6.1.12->ipykernel>=4.5.1->ipywidgets>=7.2.1->pythreejs) (6.8.0) - Requirement already satisfied: python-dateutil>=2.8.2 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from jupyter-client>=6.1.12->ipykernel>=4.5.1->ipywidgets>=7.2.1->pythreejs) (2.8.2) - Requirement already satisfied: platformdirs>=2.5 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from jupyter-core!=5.0.*,>=4.12->ipykernel>=4.5.1->ipywidgets>=7.2.1->pythreejs) (3.8.1) - Requirement already satisfied: ptyprocess>=0.5 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from pexpect>4.3->ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (0.7.0) - Requirement already satisfied: wcwidth in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from prompt-toolkit!=3.0.37,<3.1.0,>=3.0.30->ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (0.2.6) - Requirement already satisfied: executing>=1.2.0 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from stack-data->ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (1.2.0) - Requirement already satisfied: asttokens>=2.1.0 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from stack-data->ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (2.2.1) - Requirement already satisfied: pure-eval in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from stack-data->ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (0.2.2) - Requirement already satisfied: six in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from asttokens>=2.1.0->stack-data->ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (1.16.0) - Requirement already satisfied: zipp>=0.5 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from importlib-metadata>=4.8.3->jupyter-client>=6.1.12->ipykernel>=4.5.1->ipywidgets>=7.2.1->pythreejs) (3.16.0) + Requirement already satisfied: comm>=0.1.3 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipywidgets>=7.2.1->pythreejs) (0.1.4) + Requirement already satisfied: ipython>=6.1.0 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipywidgets>=7.2.1->pythreejs) (8.12.2) + Requirement already satisfied: widgetsnbextension~=4.0.7 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipywidgets>=7.2.1->pythreejs) (4.0.8) + Requirement already satisfied: jupyterlab-widgets~=3.0.7 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipywidgets>=7.2.1->pythreejs) (3.0.8) + Requirement already satisfied: backcall in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (0.2.0) + Requirement already satisfied: decorator in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (4.4.2) + Requirement already satisfied: jedi>=0.16 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (0.19.0) + Requirement already satisfied: matplotlib-inline in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (0.1.6) + Requirement already satisfied: pickleshare in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (0.7.5) + Requirement already satisfied: prompt-toolkit!=3.0.37,<3.1.0,>=3.0.30 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (3.0.39) + Requirement already satisfied: pygments>=2.4.0 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (2.16.1) + Requirement already satisfied: stack-data in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (0.6.2) + Requirement already satisfied: typing-extensions in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (4.7.1) + Requirement already satisfied: pexpect>4.3 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (4.8.0) + Requirement already satisfied: parso<0.9.0,>=0.8.3 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from jedi>=0.16->ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (0.8.3) + Requirement already satisfied: ptyprocess>=0.5 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from pexpect>4.3->ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (0.7.0) + Requirement already satisfied: wcwidth in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from prompt-toolkit!=3.0.37,<3.1.0,>=3.0.30->ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (0.2.6) + Requirement already satisfied: executing>=1.2.0 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from stack-data->ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (1.2.0) + Requirement already satisfied: asttokens>=2.1.0 in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from stack-data->ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (2.2.1) + Requirement already satisfied: pure-eval in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from stack-data->ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (0.2.2) + Requirement already satisfied: six in /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages (from asttokens>=2.1.0->stack-data->ipython>=6.1.0->ipywidgets>=7.2.1->pythreejs) (1.16.0) + Using cached ipydatawidgets-4.3.5-py2.py3-none-any.whl (271 kB) + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 Installing collected packages: traittypes, ipydatawidgets, pythreejs Successfully installed ipydatawidgets-4.3.5 pythreejs-2.4.2 traittypes-0.2.1 -Imports -------- +Imports `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -108,11 +128,13 @@ Imports import engine.engine3js as engine from engine.parse_poses import parse_poses -The model ---------- +The model `⇑ <#top>`__ +############################################################################################################################### + + +Download the model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Download the model -~~~~~~~~~~~~~~~~~~ We use ``omz_downloader``, which is a command line tool from the ``openvino-dev`` package. ``omz_downloader`` automatically creates a @@ -153,13 +175,14 @@ directory structure and downloads the selected model. -Convert Model to OpenVINO IR format -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Convert Model to OpenVINO IR format `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -The selected model comes from the public directory, which means it must -be converted into OpenVINO Intermediate Representation (OpenVINO IR). We -use ``omz_converter`` to convert the ONNX format model to the OpenVINO -IR format. + The selected model +comes from the public directory, which means it must be converted into +OpenVINO Intermediate Representation (OpenVINO IR). We use +``omz_converter`` to convert the ONNX format model to the OpenVINO IR +format. .. code:: ipython3 @@ -177,23 +200,52 @@ IR format. .. parsed-literal:: ========== Converting human-pose-estimation-3d-0001 to ONNX - Conversion to ONNX command: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/bin/python -- /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/model_zoo/internal_scripts/pytorch_to_onnx.py --model-path=model/public/human-pose-estimation-3d-0001 --model-name=PoseEstimationWithMobileNet --model-param=is_convertible_by_mo=True --import-module=model --weights=model/public/human-pose-estimation-3d-0001/human-pose-estimation-3d-0001.pth --input-shape=1,3,256,448 --input-names=data --output-names=features,heatmaps,pafs --output-file=model/public/human-pose-estimation-3d-0001/human-pose-estimation-3d-0001.onnx + Conversion to ONNX command: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/bin/python -- /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/lib/python3.8/site-packages/openvino/model_zoo/internal_scripts/pytorch_to_onnx.py --model-path=model/public/human-pose-estimation-3d-0001 --model-name=PoseEstimationWithMobileNet --model-param=is_convertible_by_mo=True --import-module=model --weights=model/public/human-pose-estimation-3d-0001/human-pose-estimation-3d-0001.pth --input-shape=1,3,256,448 --input-names=data --output-names=features,heatmaps,pafs --output-file=model/public/human-pose-estimation-3d-0001/human-pose-estimation-3d-0001.onnx ONNX check passed successfully. ========== Converting human-pose-estimation-3d-0001 to IR (FP32) - Conversion command: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/bin/python -- /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/.venv/bin/mo --framework=onnx --output_dir=/tmp/tmp2rmla_sx --model_name=human-pose-estimation-3d-0001 --input=data '--mean_values=data[128.0,128.0,128.0]' '--scale_values=data[255.0,255.0,255.0]' --output=features,heatmaps,pafs --input_model=model/public/human-pose-estimation-3d-0001/human-pose-estimation-3d-0001.onnx '--layout=data(NCHW)' '--input_shape=[1, 3, 256, 448]' --compress_to_fp16=False + Conversion command: /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/bin/python -- /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/.venv/bin/mo --framework=onnx --output_dir=/tmp/tmpgwxi10io --model_name=human-pose-estimation-3d-0001 --input=data '--mean_values=data[128.0,128.0,128.0]' '--scale_values=data[255.0,255.0,255.0]' --output=features,heatmaps,pafs --input_model=model/public/human-pose-estimation-3d-0001/human-pose-estimation-3d-0001.onnx '--layout=data(NCHW)' '--input_shape=[1, 3, 256, 448]' --compress_to_fp16=False [ INFO ] The model was converted to IR v11, the latest model format that corresponds to the source DL framework input/output format. While IR v11 is backwards compatible with OpenVINO Inference Engine API v1.0, please use API v2.0 (as of 2022.1) to take advantage of the latest improvements in IR v11. - Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/latest/openvino_2_0_transition_guide.html + Find more information about API v2.0 and IR v11 at https://docs.openvino.ai/2023.0/openvino_2_0_transition_guide.html [ SUCCESS ] Generated IR version 11 model. - [ SUCCESS ] XML file: /tmp/tmp2rmla_sx/human-pose-estimation-3d-0001.xml - [ SUCCESS ] BIN file: /tmp/tmp2rmla_sx/human-pose-estimation-3d-0001.bin + [ SUCCESS ] XML file: /tmp/tmpgwxi10io/human-pose-estimation-3d-0001.xml + [ SUCCESS ] BIN file: /tmp/tmpgwxi10io/human-pose-estimation-3d-0001.bin -Load the model -~~~~~~~~~~~~~~ +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + core = Core() + + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) + + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +Load the model `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Converted models are located in a fixed structure, which indicates vendor, model name and precision. @@ -209,8 +261,8 @@ created to infer the compiled model. ie_core = Core() # read the network and corresponding weights from file model = ie_core.read_model(model=ir_model_path, weights=model_weights_path) - # load the model on the CPU (you can also use GPU) - compiled_model = ie_core.compile_model(model=model, device_name="CPU") + # load the model on the specified device + compiled_model = ie_core.compile_model(model=model, device_name=device.value) infer_request = compiled_model.create_infer_request() input_tensor_name = model.inputs[0].get_any_name() @@ -234,15 +286,15 @@ heat maps, PAF (part affinity fields) and features. -Processing ----------- +Processing `⇑ <#top>`__ +############################################################################################################################### -Model Inference -~~~~~~~~~~~~~~~ +Model Inference `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Frames captured from video files or the live webcam are used as the -input for the 3D model. This is how you obtain the output heat maps, PAF -(part affinity fields) and features. +Frames captured from video files or the live webcam are used as the input for the 3D +model. This is how you obtain the output heat maps, PAF (part affinity +fields) and features. .. code:: ipython3 @@ -275,14 +327,13 @@ input for the 3D model. This is how you obtain the output heat maps, PAF return results -Draw 2D Pose Overlays -~~~~~~~~~~~~~~~~~~~~~ +Draw 2D Pose Overlays `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -We need to define some connections between the joints in advance, so -that we can draw the structure of the human body in the resulting image -after obtaining the inference results. Joints are drawn as circles and -limbs are drawn as lines. The code is based on the `3D Human Pose -Estimation +We need to define some connections between the joints in advance, so that we can draw the structure of the +human body in the resulting image after obtaining the inference results. +Joints are drawn as circles and limbs are drawn as lines. The code is +based on the `3D Human Pose Estimation Demo `__ from Open Model Zoo. @@ -357,8 +408,9 @@ from Open Model Zoo. return frame -Main Processing Function -~~~~~~~~~~~~~~~~~~~~~~~~ +Main Processing Function `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + Run 3D pose estimation on the specified source. It could be either a webcam feed or a video file. @@ -521,11 +573,13 @@ webcam feed or a video file. if skeleton_set: engine3D.scene_remove(skeleton_set) -Run ---- +Run `⇑ <#top>`__ +############################################################################################################################### + + +Run Live Pose Estimation `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Run Live Pose Estimation -~~~~~~~~~~~~~~~~~~~~~~~~ Run, using a webcam as the video input. By default, the primary webcam is set with ``source=0``. If you have multiple webcams, each one will be @@ -550,8 +604,9 @@ picture on the left to interact. run_pose_estimation(source=0, flip=True, use_popup=False) -Run Pose Estimation on a Video File -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Run Pose Estimation on a Video File `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + If you do not have a webcam, you can still run this demo with a video file. Any `format supported by diff --git a/docs/notebooks/407-person-tracking-with-output.rst b/docs/notebooks/407-person-tracking-with-output.rst index 0255e15626d..99aa63a8f16 100644 --- a/docs/notebooks/407-person-tracking-with-output.rst +++ b/docs/notebooks/407-person-tracking-with-output.rst @@ -1,6 +1,8 @@ Person Tracking with OpenVINO™ ============================== +.. _top: + This notebook demonstrates live person tracking with OpenVINO: it reads frames from an input video sequence, detects people in the frames, uniquely identifies each one of them and tracks all of them until they @@ -71,9 +73,9 @@ made of three key components which are as follows: |deepsort| union(IOU) association as proposed in the original SORT algorithm [3] on the set of unconfirmed and unmatched tracks from the previous step. If the IOU of detection and target is less than a certain - threshold value called IOUmin then that assignment is rejected. This - helps to account for sudden appearance changes, for example, due to - partial occlusion with static scene geometry, and to increase + threshold value called ``IOUmin`` then that assignment is rejected. + This helps to account for sudden appearance changes, for example, due + to partial occlusion with static scene geometry, and to increase robustness against erroneous. When detection result is associated with a target, the detected @@ -86,20 +88,49 @@ Problems”, Journal of Basic Engineering, vol. 82, no. Series D, pp. 35-45, 1960. [2] H. W. Kuhn, “The Hungarian method for the assignment problem”, Naval -ResearchLogistics Quarterly, vol. 2, pp. 83-97, 1955. +Research Logistics Quarterly, vol. 2, pp. 83-97, 1955. [3] A. Bewley, G. Zongyuan, F. Ramos, and B. Upcroft, “Simple online and realtime tracking,” in ICIP, 2016, pp. 3464–3468. .. |deepsort| image:: https://user-images.githubusercontent.com/91237924/221744683-0042eff8-2c41-43b8-b3ad-b5929bafb60b.png +**Table of contents**: + +- `Imports <#imports>`__ +- `Download the Model <#download-the-model>`__ +- `Load model <#load-model>`__ + + - `Select inference device <#select-inference-device>`__ + +- `Data Processing <#data-processing>`__ +- `Test person reidentification model <#test-person-reidentification-model>`__ + + - `Visualize data <#visualize-data>`__ + - `Compare two persons <#compare-two-persons>`__ + +- `Main Processing Function <#main-processing-function>`__ +- `Run <#run>`__ + + - `Initialize tracker <#initialize-tracker>`__ + - `Run Live Person Tracking <#run-live-person-tracking>`__ + - `Run Person Tracking on a Video File <#run-person-tracking-on-a-video-file>`__ + .. code:: ipython3 - !pip install -q 'openvino-dev>=2023.0.0' + !pip install -q "openvino-dev>=2023.0.0" !pip install -q opencv-python matplotlib requests scipy -Imports -------- + +.. parsed-literal:: + + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + + +Imports `⇑ <#top>`__ +############################################################################################################################### + .. code:: ipython3 @@ -134,12 +165,11 @@ Imports from deepsort_utils.nn_matching import NearestNeighborDistanceMetric from deepsort_utils.detection import Detection, compute_color_for_labels, xywh_to_xyxy, xywh_to_tlwh, tlwh_to_xyxy -Download the Model ------------------- +Download the Model `⇑ <#top>`__ +############################################################################################################################### -We will use pre-trained models from OpenVINO’s `Open Model -Zoo `__ to start the -test. +We will use pre-trained models from OpenVINO’s `Open Model Zoo `__ +to start the test. Use ``omz_downloader``, which is a command-line tool from the ``openvino-dev`` package. It automatically creates a directory structure @@ -148,22 +178,19 @@ already downloaded. The selected model comes from the public directory, which means it must be converted into OpenVINO Intermediate Representation (OpenVINO IR). - **NOTE**: Using a model outside the list can require different pre- - and post-processing. +.. note:: -In this case, `person detection -model `__ + Using a model outside the list can require different pre- and post-processing. + +In this case, `person detection model `__ is deployed to detect the person in each frame of the video, and -`reidentification -model `__ +`reidentification model `__ is used to output embedding vector to match a pair of images of a person by the cosine distance. If you want to download another model (``person-detection-xxx`` from -`Object Detection Models -list `__, -``person-reidentification-retail-xxx`` from `Reidentification Models -list `__), +`Object Detection Models list `__, +``person-reidentification-retail-xxx`` from `Reidentification Models list `__), replace the name of the model in the code below. .. code:: ipython3 @@ -219,8 +246,8 @@ replace the name of the model in the code below. -Load model ----------- +Load model `⇑ <#top>`__ +############################################################################################################################### Define a common class for model loading and predicting. @@ -238,7 +265,7 @@ performance, but slightly longer startup time). .. code:: ipython3 - ie_core = Core() + core = Core() class Model: @@ -256,7 +283,7 @@ performance, but slightly longer startup time). batchsize: batch size of input data device: device used to run inference """ - self.model = ie_core.read_model(model=model_path) + self.model = core.read_model(model=model_path) self.input_layer = self.model.input(0) self.input_shape = self.input_layer.shape self.height = self.input_shape[2] @@ -266,7 +293,7 @@ performance, but slightly longer startup time). input_shape = layer.partial_shape input_shape[0] = batchsize self.model.reshape({layer: input_shape}) - self.compiled_model = ie_core.compile_model(model=self.model, device_name=device) + self.compiled_model = core.compile_model(model=self.model, device_name=device) self.output_layer = self.compiled_model.output(0) def predict(self, input): @@ -279,20 +306,48 @@ performance, but slightly longer startup time). """ result = self.compiled_model(input)[self.output_layer] return result + +Select inference device `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + + +Select device from dropdown list for running inference using OpenVINO: + +.. code:: ipython3 + + import ipywidgets as widgets + device = widgets.Dropdown( + options=core.available_devices + ["AUTO"], + value='AUTO', + description='Device:', + disabled=False, + ) - detector = Model(detection_model_path) + device + + + + +.. parsed-literal:: + + Dropdown(description='Device:', index=1, options=('CPU', 'AUTO'), value='AUTO') + + + +.. code:: ipython3 + + detector = Model(detection_model_path, device=device.value) # since the number of detection object is uncertain, the input batch size of reid model should be dynamic - extractor = Model(reidentification_model_path, -1) + extractor = Model(reidentification_model_path, -1, device.value) -Data Processing ---------------- +Data Processing `⇑ <#top>`__ +############################################################################################################################### -Data Processing includes data preprocess and postprocess functions. - -Data preprocess function is used to change the layout and shape of input -data, according to requirement of the network input format. - Data -postprocess function is used to extract the useful information from -network’s original output and visualize it. +Data Processing includes data preprocess and postprocess functions. - Data preprocess function is used to change +the layout and shape of input data, according to requirement of the +network input format. - Data postprocess function is used to extract the +useful information from network’s original output and visualize it. .. code:: ipython3 @@ -404,15 +459,16 @@ network’s original output and visualize it. """ return np.dot(x1, x2) / (np.linalg.norm(x1) * np.linalg.norm(x2)) -Test person reidentification model ----------------------------------- +Test person reidentification model `⇑ <#top>`__ +############################################################################################################################### -The reidentification network outputs a blob with the ``(1, 256)`` shape -named ``reid_embedding``, which can be compared with other descriptors -using the cosine distance. +The reidentification network outputs a blob with the ``(1, 256)`` shape named +``reid_embedding``, which can be compared with other descriptors using +the cosine distance. + +Visualize data `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Visualize data -~~~~~~~~~~~~~~ .. code:: ipython3 @@ -456,11 +512,12 @@ Visualize data -.. image:: 407-person-tracking-with-output_files/407-person-tracking-with-output_13_3.png +.. image:: 407-person-tracking-with-output_files/407-person-tracking-with-output_17_3.png -Compare two persons -~~~~~~~~~~~~~~~~~~~ +Compare two persons `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + .. code:: ipython3 @@ -481,8 +538,9 @@ Compare two persons Different person (confidence: 0.02726622298359871) -Main Processing Function ------------------------- +Main Processing Function `⇑ <#top>`__ +############################################################################################################################### + Run person tracking on the specified source. Either a webcam feed or a video file. @@ -636,14 +694,14 @@ video file. if use_popup: cv2.destroyAllWindows() -Run ---- +Run `⇑ <#top>`__ +############################################################################################################################### -Initialize tracker -~~~~~~~~~~~~~~~~~~ -Before running a new tracking task, we have to reinitialize a Tracker -object +Initialize tracker `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + +Before running a new tracking task, we have to reinitialize a Tracker object .. code:: ipython3 @@ -659,15 +717,14 @@ object n_init=3 ) -Run Live Person Tracking -~~~~~~~~~~~~~~~~~~~~~~~~ +Run Live Person Tracking `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ -Use a webcam as the video input. By default, the primary webcam is set -with ``source=0``. If you have multiple webcams, each one will be -assigned a consecutive number starting at 0. Set ``flip=True`` when -using a front-facing camera. Some web browsers, especially Mozilla -Firefox, may cause flickering. If you experience flickering, set -``use_popup=True``. +Use a webcam as the video input. By default, the primary webcam is set with ``source=0``. If you have +multiple webcams, each one will be assigned a consecutive number +starting at 0. Set ``flip=True`` when using a front-facing camera. Some +web browsers, especially Mozilla Firefox, may cause flickering. If you +experience flickering, set ``use_popup=True``. .. code:: ipython3 @@ -681,16 +738,16 @@ Firefox, may cause flickering. If you experience flickering, set .. parsed-literal:: - [ WARN:0@10.093] global cap_v4l.cpp:982 open VIDEOIO(V4L2:/dev/video0): can't open camera by index - [ERROR:0@10.094] global obsensor_uvc_stream_channel.cpp:156 getStreamChannelGroup Camera index out of range + [ WARN:0@10.127] global cap_v4l.cpp:982 open VIDEOIO(V4L2:/dev/video0): can't open camera by index + [ERROR:0@10.127] global obsensor_uvc_stream_channel.cpp:156 getStreamChannelGroup Camera index out of range -Run Person Tracking on a Video File -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Run Person Tracking on a Video File `⇑ <#top>`__ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + If you do not have a webcam, you can still run this demo with a video -file. Any `format supported by -OpenCV `__ +file. Any `format supported by OpenCV `__ will work. .. code:: ipython3 @@ -700,7 +757,7 @@ will work. -.. image:: 407-person-tracking-with-output_files/407-person-tracking-with-output_23_0.png +.. image:: 407-person-tracking-with-output_files/407-person-tracking-with-output_27_0.png .. parsed-literal:: diff --git a/docs/notebooks/407-person-tracking-with-output_files/407-person-tracking-with-output_13_3.png b/docs/notebooks/407-person-tracking-with-output_files/407-person-tracking-with-output_17_3.png similarity index 100% rename from docs/notebooks/407-person-tracking-with-output_files/407-person-tracking-with-output_13_3.png rename to docs/notebooks/407-person-tracking-with-output_files/407-person-tracking-with-output_17_3.png diff --git a/docs/notebooks/407-person-tracking-with-output_files/407-person-tracking-with-output_23_0.png b/docs/notebooks/407-person-tracking-with-output_files/407-person-tracking-with-output_23_0.png deleted file mode 100644 index 73a5c51307f..00000000000 --- a/docs/notebooks/407-person-tracking-with-output_files/407-person-tracking-with-output_23_0.png +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:bf322a2e710ca3098dd974bccaa90b34fdefe925475886d3a24fe4fc6be30df9 -size 219400 diff --git a/docs/notebooks/407-person-tracking-with-output_files/407-person-tracking-with-output_27_0.png b/docs/notebooks/407-person-tracking-with-output_files/407-person-tracking-with-output_27_0.png new file mode 100644 index 00000000000..019e835e5d8 --- /dev/null +++ b/docs/notebooks/407-person-tracking-with-output_files/407-person-tracking-with-output_27_0.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:d9516ecd60d4a653c7f96eb72800e085c6ba50f45c8f192b41d2e431e38e2678 +size 218848 diff --git a/docs/notebooks/407-person-tracking-with-output_files/index.html b/docs/notebooks/407-person-tracking-with-output_files/index.html index 3a884de7d62..905ad643820 100644 --- a/docs/notebooks/407-person-tracking-with-output_files/index.html +++ b/docs/notebooks/407-person-tracking-with-output_files/index.html @@ -1,8 +1,8 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/407-person-tracking-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/407-person-tracking-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/407-person-tracking-with-output_files/


../
-407-person-tracking-with-output_13_3.png           12-Jul-2023 00:11              106259
-407-person-tracking-with-output_23_0.png           12-Jul-2023 00:11              219400
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/407-person-tracking-with-output_files/


../
+407-person-tracking-with-output_17_3.png           16-Aug-2023 01:31              106259
+407-person-tracking-with-output_27_0.png           16-Aug-2023 01:31              218848
 

diff --git a/docs/notebooks/index.html b/docs/notebooks/index.html index ed1a20ae3af..cd9a21671a4 100644 --- a/docs/notebooks/index.html +++ b/docs/notebooks/index.html @@ -1,154 +1,166 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/


../
-001-hello-world-with-output_files/                 12-Jul-2023 00:11                   -
-003-hello-segmentation-with-output_files/          12-Jul-2023 00:11                   -
-004-hello-detection-with-output_files/             12-Jul-2023 00:11                   -
-101-tensorflow-classification-to-openvino-with-..> 12-Jul-2023 00:11                   -
-102-pytorch-onnx-to-openvino-with-output_files/    12-Jul-2023 00:11                   -
-102-pytorch-to-openvino-with-output_files/         12-Jul-2023 00:11                   -
-103-paddle-to-openvino-classification-with-outp..> 12-Jul-2023 00:11                   -
-106-auto-device-with-output_files/                 12-Jul-2023 00:11                   -
-109-latency-tricks-with-output_files/              12-Jul-2023 00:11                   -
-110-ct-scan-live-inference-with-output_files/      12-Jul-2023 00:11                   -
-110-ct-segmentation-quantize-nncf-with-output_f..> 12-Jul-2023 00:11                   -
-111-yolov5-quantization-migration-with-output_f..> 12-Jul-2023 00:11                   -
-113-image-classification-quantization-with-outp..> 12-Jul-2023 00:11                   -
-115-async-api-with-output_files/                   12-Jul-2023 00:11                   -
-117-model-server-with-output_files/                12-Jul-2023 00:11                   -
-118-optimize-preprocessing-with-output_files/      12-Jul-2023 00:11                   -
-119-tflite-to-openvino-with-output_files/          12-Jul-2023 00:11                   -
-120-tensorflow-object-detection-to-openvino-wit..> 12-Jul-2023 00:11                   -
-201-vision-monodepth-with-output_files/            12-Jul-2023 00:11                   -
-202-vision-superresolution-image-with-output_files 12-Jul-2023 00:11                   -
-203-meter-reader-with-output_files/                12-Jul-2023 00:11                   -
-204-segmenter-semantic-segmentation-with-output..> 12-Jul-2023 00:11                   -
-205-vision-background-removal-with-output_files/   12-Jul-2023 00:11                   -
-206-vision-paddlegan-anime-with-output_files/      12-Jul-2023 00:11                   -
-207-vision-paddlegan-superresolution-with-outpu..> 12-Jul-2023 00:11                   -
-208-optical-character-recognition-with-output_f..> 12-Jul-2023 00:11                   -
-209-handwritten-ocr-with-output_files/             12-Jul-2023 00:11                   -
-211-speech-to-text-with-output_files/              12-Jul-2023 00:11                   -
-212-pyannote-speaker-diarization-with-output_files 12-Jul-2023 00:11                   -
-215-image-inpainting-with-output_files/            12-Jul-2023 00:11                   -
-216-attention-center-with-output_files/            12-Jul-2023 00:11                   -
-217-vision-deblur-with-output_files/               12-Jul-2023 00:11                   -
-218-vehicle-detection-and-recognition-with-outp..> 12-Jul-2023 00:11                   -
-222-vision-image-colorization-with-output_files/   12-Jul-2023 00:11                   -
-224-3D-segmentation-point-clouds-with-output_files 12-Jul-2023 00:11                   -
-225-stable-diffusion-text-to-image-with-output_..> 12-Jul-2023 00:11                   -
-226-yolov7-optimization-with-output_files/         12-Jul-2023 00:11                   -
-228-clip-zero-shot-image-classification-with-ou..> 12-Jul-2023 00:11                   -
-230-yolov8-optimization-with-output_files/         12-Jul-2023 00:11                   -
-231-instruct-pix2pix-image-editing-with-output_..> 12-Jul-2023 00:11                   -
-232-clip-language-saliency-map-with-output_files/  12-Jul-2023 00:11                   -
-233-blip-visual-language-processing-with-output..> 12-Jul-2023 00:11                   -
-234-encodec-audio-compression-with-output_files/   12-Jul-2023 00:11                   -
-235-controlnet-stable-diffusion-with-output_files/ 12-Jul-2023 00:11                   -
-236-stable-diffusion-v2-optimum-demo-comparison..> 12-Jul-2023 00:11                   -
-236-stable-diffusion-v2-optimum-demo-with-outpu..> 12-Jul-2023 00:11                   -
-236-stable-diffusion-v2-text-to-image-demo-with..> 12-Jul-2023 00:11                   -
-237-segment-anything-with-output_files/            12-Jul-2023 00:11                   -
-238-deep-floyd-if-with-output_files/               12-Jul-2023 00:11                   -
-239-image-bind-with-output_files/                  12-Jul-2023 00:11                   -
-241-riffusion-text-to-music-with-output_files/     12-Jul-2023 00:11                   -
-243-tflite-selfie-segmentation-with-output_files/  12-Jul-2023 00:11                   -
-301-tensorflow-training-openvino-nncf-with-outp..> 12-Jul-2023 00:11                   -
-301-tensorflow-training-openvino-with-output_files 12-Jul-2023 00:11                   -
-305-tensorflow-quantization-aware-training-with..> 12-Jul-2023 00:11                   -
-401-object-detection-with-output_files/            12-Jul-2023 00:11                   -
-402-pose-estimation-with-output_files/             12-Jul-2023 00:11                   -
-403-action-recognition-webcam-with-output_files/   12-Jul-2023 00:11                   -
-404-style-transfer-with-output_files/              12-Jul-2023 00:11                   -
-405-paddle-ocr-webcam-with-output_files/           12-Jul-2023 00:11                   -
-407-person-tracking-with-output_files/             12-Jul-2023 00:11                   -
-notebook_utils-with-output_files/                  12-Jul-2023 00:11                   -
-001-hello-world-with-output.rst                    12-Jul-2023 00:11                3706
-002-openvino-api-with-output.rst                   12-Jul-2023 00:11               32092
-003-hello-segmentation-with-output.rst             12-Jul-2023 00:11                5370
-004-hello-detection-with-output.rst                12-Jul-2023 00:11                6640
-101-tensorflow-classification-to-openvino-with-..> 12-Jul-2023 00:11                8482
-102-pytorch-onnx-to-openvino-with-output.rst       12-Jul-2023 00:11               17233
-102-pytorch-to-openvino-with-output.rst            12-Jul-2023 00:11               21693
-103-paddle-to-openvino-classification-with-outp..> 12-Jul-2023 00:11               14728
-104-model-tools-with-output.rst                    12-Jul-2023 00:11               20614
-105-language-quantize-bert-with-output.rst         12-Jul-2023 00:11               27043
-106-auto-device-with-output.rst                    12-Jul-2023 00:11               20613
-107-speech-recognition-quantization-data2vec-wi..> 12-Jul-2023 00:11              537549
-108-gpu-device-with-output.rst                     12-Jul-2023 00:11               52538
-109-latency-tricks-with-output.rst                 12-Jul-2023 00:11               21725
-110-ct-scan-live-inference-with-output.rst         12-Jul-2023 00:11               15709
-110-ct-segmentation-quantize-nncf-with-output.rst  12-Jul-2023 00:11               35493
-111-yolov5-quantization-migration-with-output.rst  12-Jul-2023 00:11               46003
-112-pytorch-post-training-quantization-nncf-wit..> 12-Jul-2023 00:11               28421
-113-image-classification-quantization-with-outp..> 12-Jul-2023 00:11               19837
-115-async-api-with-output.rst                      12-Jul-2023 00:11               17911
-116-sparsity-optimization-with-output.rst          12-Jul-2023 00:11               16707
-117-model-server-with-output.rst                   12-Jul-2023 00:11               20297
-118-optimize-preprocessing-with-output.rst         12-Jul-2023 00:11               22167
-119-tflite-to-openvino-with-output.rst             12-Jul-2023 00:11               10299
-120-tensorflow-object-detection-to-openvino-wit..> 12-Jul-2023 00:11               26120
-201-vision-monodepth-with-output.rst               12-Jul-2023 00:11              967091
-202-vision-superresolution-image-with-output.rst   12-Jul-2023 00:11               24505
-203-meter-reader-with-output.rst                   12-Jul-2023 00:11               25102
-204-segmenter-semantic-segmentation-with-output..> 12-Jul-2023 00:11               28025
-205-vision-background-removal-with-output.rst      12-Jul-2023 00:11               13925
-206-vision-paddlegan-anime-with-output.rst         12-Jul-2023 00:11               20078
-207-vision-paddlegan-superresolution-with-outpu..> 12-Jul-2023 00:11               14105
-208-optical-character-recognition-with-output.rst  12-Jul-2023 00:11               24996
-209-handwritten-ocr-with-output.rst                12-Jul-2023 00:11               10374
-210-slowfast-video-recognition-with-output.rst     12-Jul-2023 00:11              781777
-211-speech-to-text-with-output.rst                 12-Jul-2023 00:11               86228
-212-pyannote-speaker-diarization-with-output.rst   12-Jul-2023 00:11             1292382
-213-question-answering-with-output.rst             12-Jul-2023 00:11               19575
-214-grammar-correction-with-output.rst             12-Jul-2023 00:11               16212
-215-image-inpainting-with-output.rst               12-Jul-2023 00:11                8112
-216-attention-center-with-output.rst               12-Jul-2023 00:11               10427
-217-vision-deblur-with-output.rst                  12-Jul-2023 00:11                9534
-218-vehicle-detection-and-recognition-with-outp..> 12-Jul-2023 00:11               16002
-219-knowledge-graphs-conve-with-output.rst         12-Jul-2023 00:11               21692
-221-machine-translation-with-output.rst            12-Jul-2023 00:11                8095
-222-vision-image-colorization-with-output.rst      12-Jul-2023 00:11               12717
-223-text-prediction-with-output.rst                12-Jul-2023 00:11               24192
-224-3D-segmentation-point-clouds-with-output.rst   12-Jul-2023 00:11                7616
-225-stable-diffusion-text-to-image-with-output.rst 12-Jul-2023 00:11               41516
-226-yolov7-optimization-with-output.rst            12-Jul-2023 00:11               39571
-227-whisper-subtitles-generation-with-output.rst   12-Jul-2023 00:11               35931
-228-clip-zero-shot-image-classification-with-ou..> 12-Jul-2023 00:11               15734
-229-distilbert-sequence-classification-with-out..> 12-Jul-2023 00:11                6263
-230-yolov8-optimization-with-output.rst            12-Jul-2023 00:11               70077
-231-instruct-pix2pix-image-editing-with-output.rst 12-Jul-2023 00:11               46622
-232-clip-language-saliency-map-with-output.rst     12-Jul-2023 00:11               33027
-233-blip-visual-language-processing-with-output..> 12-Jul-2023 00:11               43807
-234-encodec-audio-compression-with-output.rst      12-Jul-2023 00:11             3861302
-235-controlnet-stable-diffusion-with-output.rst    12-Jul-2023 00:11               50491
-236-stable-diffusion-v2-infinite-zoom-with-outp..> 12-Jul-2023 00:11               60547
-236-stable-diffusion-v2-optimum-demo-comparison..> 12-Jul-2023 00:11                6847
-236-stable-diffusion-v2-optimum-demo-with-outpu..> 12-Jul-2023 00:11                7314
-236-stable-diffusion-v2-text-to-image-demo-with..> 12-Jul-2023 00:11               10677
-236-stable-diffusion-v2-text-to-image-with-outp..> 12-Jul-2023 00:11               53527
-237-segment-anything-with-output.rst               12-Jul-2023 00:11               59634
-238-deep-floyd-if-with-output.rst                  12-Jul-2023 00:11               28390
-239-image-bind-with-output.rst                     12-Jul-2023 00:11             2394151
-240-dolly-2-instruction-following-with-output.rst  12-Jul-2023 00:11               26853
-241-riffusion-text-to-music-with-output.rst        12-Jul-2023 00:11              621124
-242-freevc-voice-conversion-with-output.rst        12-Jul-2023 00:11              649784
-243-tflite-selfie-segmentation-with-output.rst     12-Jul-2023 00:11               20552
-301-tensorflow-training-openvino-nncf-with-outp..> 12-Jul-2023 00:11               41523
-301-tensorflow-training-openvino-with-output.rst   12-Jul-2023 00:11               40490
-302-pytorch-quantization-aware-training-with-ou..> 12-Jul-2023 00:11               29439
-305-tensorflow-quantization-aware-training-with..> 12-Jul-2023 00:11               22456
-401-object-detection-with-output.rst               12-Jul-2023 00:11               16330
-402-pose-estimation-with-output.rst                12-Jul-2023 00:11               14926
-403-action-recognition-webcam-with-output.rst      12-Jul-2023 00:11               24394
-404-style-transfer-with-output.rst                 12-Jul-2023 00:11               15389
-405-paddle-ocr-webcam-with-output.rst              12-Jul-2023 00:11               22036
-406-3D-pose-estimation-with-output.rst             12-Jul-2023 00:11               29778
-407-person-tracking-with-output.rst                12-Jul-2023 00:11               25507
-notebook_utils-with-output.rst                     12-Jul-2023 00:11               11716
-notebooks_tags.json                                12-Jul-2023 00:11                7429
-notebooks_with_binder_buttons.txt                  12-Jul-2023 00:11                 924
-notebooks_with_colab_buttons.txt                   12-Jul-2023 00:11                 698
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/


../
+001-hello-world-with-output_files/                 16-Aug-2023 01:31                   -
+003-hello-segmentation-with-output_files/          16-Aug-2023 01:31                   -
+004-hello-detection-with-output_files/             16-Aug-2023 01:31                   -
+101-tensorflow-classification-to-openvino-with-..> 16-Aug-2023 01:31                   -
+102-pytorch-onnx-to-openvino-with-output_files/    16-Aug-2023 01:31                   -
+102-pytorch-to-openvino-with-output_files/         16-Aug-2023 01:31                   -
+103-paddle-to-openvino-classification-with-outp..> 16-Aug-2023 01:31                   -
+106-auto-device-with-output_files/                 16-Aug-2023 01:31                   -
+109-latency-tricks-with-output_files/              16-Aug-2023 01:31                   -
+109-throughput-tricks-with-output_files/           16-Aug-2023 01:31                   -
+110-ct-scan-live-inference-with-output_files/      16-Aug-2023 01:31                   -
+110-ct-segmentation-quantize-nncf-with-output_f..> 16-Aug-2023 01:31                   -
+111-yolov5-quantization-migration-with-output_f..> 16-Aug-2023 01:31                   -
+113-image-classification-quantization-with-outp..> 16-Aug-2023 01:31                   -
+115-async-api-with-output_files/                   16-Aug-2023 01:31                   -
+117-model-server-with-output_files/                16-Aug-2023 01:31                   -
+118-optimize-preprocessing-with-output_files/      16-Aug-2023 01:31                   -
+119-tflite-to-openvino-with-output_files/          16-Aug-2023 01:31                   -
+120-tensorflow-object-detection-to-openvino-wit..> 16-Aug-2023 01:31                   -
+201-vision-monodepth-with-output_files/            16-Aug-2023 01:31                   -
+202-vision-superresolution-image-with-output_files 16-Aug-2023 01:31                   -
+203-meter-reader-with-output_files/                16-Aug-2023 01:31                   -
+204-segmenter-semantic-segmentation-with-output..> 16-Aug-2023 01:31                   -
+205-vision-background-removal-with-output_files/   16-Aug-2023 01:31                   -
+206-vision-paddlegan-anime-with-output_files/      16-Aug-2023 01:31                   -
+207-vision-paddlegan-superresolution-with-outpu..> 16-Aug-2023 01:31                   -
+208-optical-character-recognition-with-output_f..> 16-Aug-2023 01:31                   -
+209-handwritten-ocr-with-output_files/             16-Aug-2023 01:31                   -
+211-speech-to-text-with-output_files/              16-Aug-2023 01:31                   -
+212-pyannote-speaker-diarization-with-output_files 16-Aug-2023 01:31                   -
+215-image-inpainting-with-output_files/            16-Aug-2023 01:31                   -
+216-attention-center-with-output_files/            16-Aug-2023 01:31                   -
+217-vision-deblur-with-output_files/               16-Aug-2023 01:31                   -
+218-vehicle-detection-and-recognition-with-outp..> 16-Aug-2023 01:31                   -
+220-cross-lingual-books-alignment-with-output_f..> 16-Aug-2023 01:31                   -
+222-vision-image-colorization-with-output_files/   16-Aug-2023 01:31                   -
+224-3D-segmentation-point-clouds-with-output_files 16-Aug-2023 01:31                   -
+225-stable-diffusion-text-to-image-with-output_..> 16-Aug-2023 01:31                   -
+226-yolov7-optimization-with-output_files/         16-Aug-2023 01:31                   -
+228-clip-zero-shot-convert-with-output_files/      16-Aug-2023 01:31                   -
+228-clip-zero-shot-quantize-with-output_files/     16-Aug-2023 01:31                   -
+230-yolov8-optimization-with-output_files/         16-Aug-2023 01:31                   -
+231-instruct-pix2pix-image-editing-with-output_..> 16-Aug-2023 01:31                   -
+233-blip-visual-language-processing-with-output..> 16-Aug-2023 01:31                   -
+234-encodec-audio-compression-with-output_files/   16-Aug-2023 01:31                   -
+235-controlnet-stable-diffusion-with-output_files/ 16-Aug-2023 01:31                   -
+236-stable-diffusion-v2-optimum-demo-comparison..> 16-Aug-2023 01:31                   -
+236-stable-diffusion-v2-optimum-demo-with-outpu..> 16-Aug-2023 01:31                   -
+236-stable-diffusion-v2-text-to-image-demo-with..> 16-Aug-2023 01:31                   -
+237-segment-anything-with-output_files/            16-Aug-2023 01:31                   -
+238-deep-floyd-if-with-output_files/               16-Aug-2023 01:31                   -
+239-image-bind-convert-with-output_files/          16-Aug-2023 01:31                   -
+241-riffusion-text-to-music-with-output_files/     16-Aug-2023 01:31                   -
+243-tflite-selfie-segmentation-with-output_files/  16-Aug-2023 01:31                   -
+246-depth-estimation-videpth-with-output_files/    16-Aug-2023 01:31                   -
+248-stable-diffusion-xl-with-output_files/         16-Aug-2023 01:31                   -
+249-oneformer-segmentation-with-output_files/      16-Aug-2023 01:31                   -
+301-tensorflow-training-openvino-with-output_files 16-Aug-2023 01:31                   -
+305-tensorflow-quantization-aware-training-with..> 16-Aug-2023 01:31                   -
+401-object-detection-with-output_files/            16-Aug-2023 01:31                   -
+402-pose-estimation-with-output_files/             16-Aug-2023 01:31                   -
+403-action-recognition-webcam-with-output_files/   16-Aug-2023 01:31                   -
+404-style-transfer-with-output_files/              16-Aug-2023 01:31                   -
+405-paddle-ocr-webcam-with-output_files/           16-Aug-2023 01:31                   -
+407-person-tracking-with-output_files/             16-Aug-2023 01:31                   -
+notebook_utils-with-output_files/                  16-Aug-2023 01:31                   -
+001-hello-world-with-output.rst                    16-Aug-2023 01:31                3886
+002-openvino-api-with-output.rst                   16-Aug-2023 01:31               32160
+003-hello-segmentation-with-output.rst             16-Aug-2023 01:31                5603
+004-hello-detection-with-output.rst                16-Aug-2023 01:31                6847
+101-tensorflow-classification-to-openvino-with-..> 16-Aug-2023 01:31                8858
+102-pytorch-onnx-to-openvino-with-output.rst       16-Aug-2023 01:31               17662
+102-pytorch-to-openvino-with-output.rst            16-Aug-2023 01:31               22399
+103-paddle-to-openvino-classification-with-outp..> 16-Aug-2023 01:31               15022
+104-model-tools-with-output.rst                    16-Aug-2023 01:31               20907
+105-language-quantize-bert-with-output.rst         16-Aug-2023 01:31               27178
+106-auto-device-with-output.rst                    16-Aug-2023 01:31               21069
+107-speech-recognition-quantization-data2vec-wi..> 16-Aug-2023 01:31              537722
+108-gpu-device-with-output.rst                     16-Aug-2023 01:31               53377
+109-latency-tricks-with-output.rst                 16-Aug-2023 01:31               22521
+109-throughput-tricks-with-output.rst              16-Aug-2023 01:31               25678
+110-ct-scan-live-inference-with-output.rst         16-Aug-2023 01:31               16024
+110-ct-segmentation-quantize-nncf-with-output.rst  16-Aug-2023 01:31               35892
+111-yolov5-quantization-migration-with-output.rst  16-Aug-2023 01:31               46306
+112-pytorch-post-training-quantization-nncf-wit..> 16-Aug-2023 01:31               28796
+113-image-classification-quantization-with-outp..> 16-Aug-2023 01:31               20160
+115-async-api-with-output.rst                      16-Aug-2023 01:31               18664
+116-sparsity-optimization-with-output.rst          16-Aug-2023 01:31               16901
+117-model-server-with-output.rst                   16-Aug-2023 01:31               20692
+118-optimize-preprocessing-with-output.rst         16-Aug-2023 01:31               22961
+119-tflite-to-openvino-with-output.rst             16-Aug-2023 01:31               10579
+120-tensorflow-object-detection-to-openvino-wit..> 16-Aug-2023 01:31               26589
+121-convert-to-openvino-with-output.rst            16-Aug-2023 01:31               83068
+201-vision-monodepth-with-output.rst               16-Aug-2023 01:31              967562
+202-vision-superresolution-image-with-output.rst   16-Aug-2023 01:31               25109
+203-meter-reader-with-output.rst                   16-Aug-2023 01:31               25544
+204-segmenter-semantic-segmentation-with-output..> 16-Aug-2023 01:31               28535
+205-vision-background-removal-with-output.rst      16-Aug-2023 01:31               14792
+206-vision-paddlegan-anime-with-output.rst         16-Aug-2023 01:31               20211
+207-vision-paddlegan-superresolution-with-outpu..> 16-Aug-2023 01:31               14872
+208-optical-character-recognition-with-output.rst  16-Aug-2023 01:31               25888
+209-handwritten-ocr-with-output.rst                16-Aug-2023 01:31               10851
+210-slowfast-video-recognition-with-output.rst     16-Aug-2023 01:31              782597
+211-speech-to-text-with-output.rst                 16-Aug-2023 01:31               87059
+212-pyannote-speaker-diarization-with-output.rst   16-Aug-2023 01:31             1293516
+213-question-answering-with-output.rst             16-Aug-2023 01:31               20417
+214-grammar-correction-with-output.rst             16-Aug-2023 01:31               19552
+215-image-inpainting-with-output.rst               16-Aug-2023 01:31                8755
+216-attention-center-with-output.rst               16-Aug-2023 01:31               11247
+217-vision-deblur-with-output.rst                  16-Aug-2023 01:31               10336
+218-vehicle-detection-and-recognition-with-outp..> 16-Aug-2023 01:31               16655
+219-knowledge-graphs-conve-with-output.rst         16-Aug-2023 01:31               20390
+220-cross-lingual-books-alignment-with-output.rst  16-Aug-2023 01:31               50713
+221-machine-translation-with-output.rst            16-Aug-2023 01:31                9632
+222-vision-image-colorization-with-output.rst      16-Aug-2023 01:31               13399
+223-text-prediction-with-output.rst                16-Aug-2023 01:31               25940
+224-3D-segmentation-point-clouds-with-output.rst   16-Aug-2023 01:31                8271
+225-stable-diffusion-text-to-image-with-output.rst 16-Aug-2023 01:31               42288
+226-yolov7-optimization-with-output.rst            16-Aug-2023 01:31               41509
+227-whisper-subtitles-generation-with-output.rst   16-Aug-2023 01:31               38576
+228-clip-zero-shot-convert-with-output.rst         16-Aug-2023 01:31               15301
+228-clip-zero-shot-quantize-with-output.rst        16-Aug-2023 01:31               11960
+229-distilbert-sequence-classification-with-out..> 16-Aug-2023 01:31                6947
+230-yolov8-optimization-with-output.rst            16-Aug-2023 01:31               75132
+231-instruct-pix2pix-image-editing-with-output.rst 16-Aug-2023 01:31               47035
+233-blip-visual-language-processing-with-output..> 16-Aug-2023 01:31               46332
+234-encodec-audio-compression-with-output.rst      16-Aug-2023 01:31             3862060
+235-controlnet-stable-diffusion-with-output.rst    16-Aug-2023 01:31               53511
+236-stable-diffusion-v2-infinite-zoom-with-outp..> 16-Aug-2023 01:31               64824
+236-stable-diffusion-v2-optimum-demo-comparison..> 16-Aug-2023 01:31                7142
+236-stable-diffusion-v2-optimum-demo-with-outpu..> 16-Aug-2023 01:31                7605
+236-stable-diffusion-v2-text-to-image-demo-with..> 16-Aug-2023 01:31               11300
+236-stable-diffusion-v2-text-to-image-with-outp..> 16-Aug-2023 01:31               44768
+237-segment-anything-with-output.rst               16-Aug-2023 01:31               97904
+238-deep-floyd-if-with-output.rst                  16-Aug-2023 01:31               29347
+239-image-bind-convert-with-output.rst             16-Aug-2023 01:31             2400458
+240-dolly-2-instruction-following-with-output.rst  16-Aug-2023 01:31               34200
+241-riffusion-text-to-music-with-output.rst        16-Aug-2023 01:31              633040
+242-freevc-voice-conversion-with-output.rst        16-Aug-2023 01:31              651968
+243-tflite-selfie-segmentation-with-output.rst     16-Aug-2023 01:31               21290
+244-named-entity-recognition-with-output.rst       16-Aug-2023 01:31               29820
+245-typo-detector-with-output.rst                  16-Aug-2023 01:31               28464
+246-depth-estimation-videpth-with-output.rst       16-Aug-2023 01:31               50251
+247-code-language-id-with-output.rst               16-Aug-2023 01:31               35934
+248-stable-diffusion-xl-with-output.rst            16-Aug-2023 01:31               21901
+249-oneformer-segmentation-with-output.rst         16-Aug-2023 01:31               14678
+301-tensorflow-training-openvino-with-output.rst   16-Aug-2023 01:31               41424
+302-pytorch-quantization-aware-training-with-ou..> 16-Aug-2023 01:31               29746
+305-tensorflow-quantization-aware-training-with..> 16-Aug-2023 01:31               23795
+401-object-detection-with-output.rst               16-Aug-2023 01:31               18234
+402-pose-estimation-with-output.rst                16-Aug-2023 01:31               15731
+403-action-recognition-webcam-with-output.rst      16-Aug-2023 01:31               25273
+404-style-transfer-with-output.rst                 16-Aug-2023 01:31               16087
+405-paddle-ocr-webcam-with-output.rst              16-Aug-2023 01:31               24119
+406-3D-pose-estimation-with-output.rst             16-Aug-2023 01:31               28035
+407-person-tracking-with-output.rst                16-Aug-2023 01:31               27202
+notebook_utils-with-output.rst                     16-Aug-2023 01:31               12520
+notebooks_tags.json                                16-Aug-2023 01:31                8554
+notebooks_with_binder_buttons.txt                  16-Aug-2023 01:31                1000
+notebooks_with_colab_buttons.txt                   16-Aug-2023 01:31                 782
 

diff --git a/docs/notebooks/notebook_utils-with-output.rst b/docs/notebooks/notebook_utils-with-output.rst index 6c632166a08..ea2ec2dffa9 100644 --- a/docs/notebooks/notebook_utils-with-output.rst +++ b/docs/notebooks/notebook_utils-with-output.rst @@ -23,6 +23,13 @@ functions in the section. !pip install -q "openvino>=2023.0.0" opencv-python !pip install -q pillow tqdm requests matplotlib + +.. parsed-literal:: + + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + DEPRECATION: pytorch-lightning 1.6.5 has a non-standard dependency specifier torch>=1.8.*. pip 23.3 will enforce this behaviour change. A possible replacement is to upgrade to a newer version of pytorch-lightning or contact the author to suggest that they release a version with a conforming dependency specifiers. Discussion can be found at https://github.com/pypa/pip/issues/12063 + + Files ----- @@ -84,7 +91,7 @@ Test File Functions .. parsed-literal:: - /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-448/.workspace/scm/ov-notebook/notebooks/utils/Safety_Full_Hat_and_Vest.mp4 + /opt/home/k8sworker/ci-ai/cibuilds/ov-notebook/OVNotebookOps-475/.workspace/scm/ov-notebook/notebooks/utils/Safety_Full_Hat_and_Vest.mp4 .. code:: ipython3 @@ -101,12 +108,12 @@ Test File Functions .. parsed-literal:: - openvino_notebooks_readme.md: 0%| | 0.00/10.2k [00:00This notebook requires OpenVINO 2022.1. The version on your system is: 2023.0.0-10926-b4452d56304-releases/2023/0.
Please run pip install --upgrade -r requirements.txt in the openvino_env environment to install this version. See the OpenVINO Notebooks README for detailed instructions +
This notebook requires OpenVINO 2022.1. The version on your system is: 2023.0.1-11005-fa1c41994f3-releases/2023/0.
Please run pip install --upgrade -r requirements.txt in the openvino_env environment to install this version. See the OpenVINO Notebooks README for detailed instructions diff --git a/docs/notebooks/notebook_utils-with-output_files/index.html b/docs/notebooks/notebook_utils-with-output_files/index.html index 54f8bfbfbb2..02c4762765e 100644 --- a/docs/notebooks/notebook_utils-with-output_files/index.html +++ b/docs/notebooks/notebook_utils-with-output_files/index.html @@ -1,13 +1,13 @@ -Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/notebook_utils-with-output_files/ +Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/notebook_utils-with-output_files/ -

Index of /projects/ov-notebook/0.1.0-latest/20230711220806/dist/rst_files/notebook_utils-with-output_files/


../
-notebook_utils-with-output_12_0.jpg                12-Jul-2023 00:11              121563
-notebook_utils-with-output_12_0.png                12-Jul-2023 00:11              869307
-notebook_utils-with-output_26_0.png                12-Jul-2023 00:11               45415
-notebook_utils-with-output_41_0.png                12-Jul-2023 00:11               10059
-notebook_utils-with-output_41_1.png                12-Jul-2023 00:11               37584
-notebook_utils-with-output_41_2.png                12-Jul-2023 00:11               16690
-notebook_utils-with-output_41_3.png                12-Jul-2023 00:11               38992
+

Index of /projects/ov-notebook/0.1.0-latest/20230815220807/dist/rst_files/notebook_utils-with-output_files/


../
+notebook_utils-with-output_12_0.jpg                16-Aug-2023 01:31              121563
+notebook_utils-with-output_12_0.png                16-Aug-2023 01:31              869307
+notebook_utils-with-output_26_0.png                16-Aug-2023 01:31               46994
+notebook_utils-with-output_41_0.png                16-Aug-2023 01:31               10059
+notebook_utils-with-output_41_1.png                16-Aug-2023 01:31               37584
+notebook_utils-with-output_41_2.png                16-Aug-2023 01:31               16690
+notebook_utils-with-output_41_3.png                16-Aug-2023 01:31               38992
 

diff --git a/docs/notebooks/notebook_utils-with-output_files/notebook_utils-with-output_26_0.png b/docs/notebooks/notebook_utils-with-output_files/notebook_utils-with-output_26_0.png index 2a532ddebda..970378fa67d 100644 --- a/docs/notebooks/notebook_utils-with-output_files/notebook_utils-with-output_26_0.png +++ b/docs/notebooks/notebook_utils-with-output_files/notebook_utils-with-output_26_0.png @@ -1,3 +1,3 @@ version https://git-lfs.github.com/spec/v1 -oid sha256:da7fb0f49e8a020cb9670630b82b4a7a76f18ecf8effd03c47520db910c716bd -size 45415 +oid sha256:3fe1fec7d917f0af31e6bee93a056bf53355c8bbff57d519a093c313d67ba1d6 +size 46994 diff --git a/docs/notebooks/notebooks_tags.json b/docs/notebooks/notebooks_tags.json index 5f88b5497ce..e6899ffc0c3 100644 --- a/docs/notebooks/notebooks_tags.json +++ b/docs/notebooks/notebooks_tags.json @@ -61,6 +61,12 @@ "ONNX", "Pytorch" ], + "109-throughput-tricks": [ + "Async Inference", + "Benchmark Model", + "ONNX", + "Pytorch" + ], "110-ct-scan-live-inference": [ "Async Inference", "Benchmark Model" @@ -106,6 +112,12 @@ "119-tflite-to-openvino": [ "Benchmark Model" ], + "121-convert-to-openvino": [ + "ONNX", + "Pytorch", + "Torchvision", + "Transformers" + ], "203-meter-reader": [ "Dynamic Shape", "Reshape Model" @@ -174,6 +186,12 @@ "ONNX", "Pytorch" ], + "220-cross-lingual-books-alignment": [ + "Async Inference", + "Optimize Model", + "Pytorch", + "Transformers" + ], "221-machine-translation": [ "Download Model" ], @@ -203,13 +221,20 @@ ], "227-whisper-subtitles-generation": [ "ONNX", - "Pytorch" + "Pytorch", + "Transformers" ], - "228-clip-zero-shot-image-classification": [ + "228-clip-zero-shot-convert": [ "ONNX", "Pytorch", "Transformers" ], + "228-clip-zero-shot-quantize": [ + "Benchmark Model", + "NNCF", + "Pytorch", + "Transformers" + ], "229-distilbert-sequence-classification": [ "Transformers" ], @@ -278,10 +303,15 @@ "Pytorch", "Reshape Model" ], - "239-image-bind": [ + "239-image-bind-convert": [ "ONNX", "Pytorch" ], + "239-image-bind-quantize": [ + "Benchmark Model", + "NNCF", + "Pytorch" + ], "240-dolly-2-instruction-following": [ "Transformers" ], @@ -292,6 +322,30 @@ "ONNX", "Pytorch" ], + "244-named-entity-recognition": [ + "ONNX", + "Pytorch", + "Transformers" + ], + "245-typo-detector": [ + "ONNX", + "Optimize Model", + "Pytorch", + "Transformers" + ], + "246-depth-estimation-videpth": [ + "ONNX", + "Pytorch", + "Torchvision" + ], + "247-code-language-id": [ + "ONNX", + "Transformers" + ], + "249-oneformer-segmentation": [ + "Pytorch", + "Transformers" + ], "301-tensorflow-training-openvino-nncf": [ "Benchmark Model", "NNCF", diff --git a/docs/notebooks/notebooks_with_binder_buttons.txt b/docs/notebooks/notebooks_with_binder_buttons.txt index 9ed3271a047..6a70fccc790 100644 --- a/docs/notebooks/notebooks_with_binder_buttons.txt +++ b/docs/notebooks/notebooks_with_binder_buttons.txt @@ -1,4 +1,3 @@ -{} 001-hello-world 002-openvino-api 003-hello-segmentation @@ -11,6 +10,7 @@ 113-image-classification-quantization 115-async-api 120-tensorflow-object-detection-to-openvino +121-convert-to-openvino 201-vision-monodepth 202-vision-superresolution-image 202-vision-superresolution-video @@ -24,10 +24,12 @@ 217-vision-deblur 218-vehicle-detection-and-recognition 219-knowledge-graphs-conve +220-cross-lingual-books-alignment 221-machine-translation 222-vision-image-colorization 229-distilbert-sequence-classification 243-tflite-selfie-segmentation +247-code-language-id 401-object-detection 402-pose-estimation 403-action-recognition-webcam diff --git a/docs/notebooks/notebooks_with_colab_buttons.txt b/docs/notebooks/notebooks_with_colab_buttons.txt index 47a8a8c7dfd..695560c6ad0 100644 --- a/docs/notebooks/notebooks_with_colab_buttons.txt +++ b/docs/notebooks/notebooks_with_colab_buttons.txt @@ -1,4 +1,3 @@ -{} 002-openvino-api 102-pytorch-to-openvino 107-speech-recognition-quantization-data2vec @@ -7,18 +6,21 @@ 116-sparsity-optimization 119-tflite-to-openvino 120-tensorflow-object-detection-to-openvino +121-convert-to-openvino 201-vision-monodepth 202-vision-superresolution-image 202-vision-superresolution-video 204-segmenter-semantic-segmentation 205-vision-background-removal 206-vision-paddlegan-anime +220-cross-lingual-books-alignment 221-machine-translation 223-text-prediction 227-whisper-subtitles-generation 230-yolov8-optimization 232-clip-language-saliency-map 243-tflite-selfie-segmentation +244-named-entity-recognition 305-tensorflow-quantization-aware-training 401-object-detection 404-style-transfer diff --git a/docs/tutorials.md b/docs/tutorials.md index edd9c91063b..f59b187c054 100644 --- a/docs/tutorials.md +++ b/docs/tutorials.md @@ -36,40 +36,28 @@ the name of it and the Jupyter notebook will start it in a new tab of a browser. on how to run and manage the notebooks on your machine. --------------------- - **Contents:** -- `Getting Started <#-getting-started>`__ +- `Getting Started <#getting-started>`__ - - `First steps with OpenVINO <#-first-steps>`__ - - `Convert & Optimize <#-convert--optimize>`__ - - `Model Demos <#-model-demos>`__ - - `Model Training <#-model-training>`__ - - `Live Demos <#-live-demos>`__ - - `Recommended Tutorials <#-recommended-tutorials>`__ - - `Additional Resources <#-additional-resources>`__ - - `Contributors <#-contributors>`__ + - `First steps with OpenVINO <#first-steps-with-openvino>`__ + - `Convert & Optimize <#convert-optimize>`__ + - `Model Demos <#model-demos>`__ + - `Model Training <#model-training>`__ + - `Live Demos <#live-demos>`__ + - `Recommended Tutorials <#recommended-tutorials>`__ + - `Additional Resources <#additional-resources>`__ + - `Contributors <#contributors>`__ --------------------- -.. raw:: html - - - -`Getting Started`_ +Getting Started ================== The Jupyter notebooks are categorized into four classes, select one related to your needs or give them all a try. Good Luck! -.. raw:: html - - - - -`First steps with OpenVINO`_ -------------------------------- +First steps with OpenVINO +------------------------- Brief tutorials that demonstrate how to use Python API for inference in OpenVINO. @@ -85,12 +73,8 @@ Brief tutorials that demonstrate how to use Python API for inference in OpenVINO | `004-hello-detection `__ |br| |n004| | Text detection with OpenVINO. | |n004-img1| | +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ -.. raw:: html - - - -`Convert & Optimize`_ ------------------------ +Convert & Optimize +-------------------- Tutorials that explain how to optimize and quantize models with OpenVINO tools. @@ -99,7 +83,7 @@ Tutorials that explain how to optimize and quantize models with OpenVINO tools. +====================================================================================================================================+============================================================================================================================================+===========================================+ | `101-tensorflow-classification-to-openvino `__ |br| |n101| | Convert TensorFlow models to OpenVINO IR. | |n101-img1| | +------------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ -| `102-pytorch-onnx-to-openvino `__ | Convert PyTorch models to OpenVINO IR. | |n102-img1| | +| `102-pytorch-to-openvino `__ |br| |c102| | Convert PyTorch models to OpenVINO IR. | |n102-img1| | +------------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ | `103-paddle-onnx-to-openvino `__ |br| |n103| | Convert PaddlePaddle models to OpenVINO IR. | |n103-img1| | +------------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ @@ -111,13 +95,17 @@ Tutorials that explain how to optimize and quantize models with OpenVINO tools. +----------------------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------+ | Notebook | Description | +====================================================================================================================================================+==================================================================================================================================+ + | `102-pytorch-onnx-to-openvino `__ | Convert PyTorch models to OpenVINO IR. | + +----------------------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------+ | `105-language-quantize-bert `__ | Optimize and quantize a pre-trained BERT model. | +----------------------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------+ | `106-auto-device `__ |br| |n106| | Demonstrates how to use AUTO Device. | +----------------------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------+ | `107-speech-recognition-quantization `__ |br| |c107| | Optimize and quantize a pre-trained Data2Vec speech model. | +----------------------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------+ - | `109-performance-tricks `__ | Performance tricks in OpenVINO™. | + | `109-latency-tricks `__ | Performance tricks for latency mode in OpenVINO™. | + +----------------------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------+ + | `109-throughput-tricks `__ | Performance tricks for throughput mode in OpenVINO™. | +----------------------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------+ | `110-ct-segmentation-quantize `__ |br| |n110| | Live inference of a kidney segmentation model and benchmark CT-scan data with OpenVINO. | +----------------------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------+ @@ -141,14 +129,12 @@ Tutorials that explain how to optimize and quantize models with OpenVINO tools. +----------------------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------+ | `120-tensorflow-object-detection-to-openvino `__ |br| |n120| |br| |c120| | Convert TensorFlow Object Detection models to OpenVINO IR | +----------------------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------+ + | `121-convert-to-openvino `__ |br| |n121| |br| |c121| | Learn OpenVINO model conversion API | + +----------------------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------+ -.. raw:: html - - - -`Model Demos`_ ----------------- +Model Demos +-------------------- Demos that demonstrate inference on a particular model. @@ -201,6 +187,8 @@ Demos that demonstrate inference on a particular model. +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ | `219-knowledge-graphs-conve `__ |br| |n219| | Optimize the knowledge graph embeddings model (ConvE) with OpenVINO. | | +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ + | `220-cross-lingual-books-alignment `__ |br| |n220| |br| |c220| | Cross-lingual Books Alignment With Transformers and OpenVINO™ | |n220-img1| | + +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ | `221-machine-translation `__ |br| |n221| |br| |c221| | Real-time translation from English to German. | | +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ | `222-vision-image-colorization `__ |br| |n222| | Use pre-trained models to colorize black & white images using OpenVINO. | |n222-img1| | @@ -215,7 +203,9 @@ Demos that demonstrate inference on a particular model. +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ | `227-whisper-subtitles-generation `__ |br| |c227| | Generate subtitles for video with OpenAI Whisper and OpenVINO. | |n227-img1| | +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ - | `228-clip-zero-shot-image-classification `__ | Perform Zero-shot image classification with CLIP and OpenVINO. | |n228-img1| | + | `228-clip-zero-shot-convert `__ | Zero-shot Image Classification with OpenAI CLIP and OpenVINO™ | |n228-img1| | + +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ + | `228-clip-zero-shot-quantize `__ | Post-Training Quantization of OpenAI CLIP model with NNCF | |n228-img2| | +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ | `229-distilbert-sequence-classification `__ |br| |n229| | Sequence classification with OpenVINO. | |n229-img1| | +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ @@ -245,7 +235,7 @@ Demos that demonstrate inference on a particular model. +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ | `238-deep-floyd-if `__ | Text-to-image generation with DeepFloyd IF and OpenVINO™. | |n238-img1| | +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ - | `239-image-bind `__ | Binding multimodal data, using ImageBind and OpenVINO™. | |n239-img1| | + | `239-image-bind `__ | Binding multimodal data, using ImageBind and OpenVINO™. | |n239-img1| | +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ | `240-dolly-2-instruction-following `__ | Instruction following using Databricks Dolly 2.0 and OpenVINO™. | |n240-img1| | +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ @@ -255,14 +245,22 @@ Demos that demonstrate inference on a particular model. +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ | `243-tflite-selfie-segmentation `__ |br| |n243| |br| |c243| | Selfie Segmentation using TFLite and OpenVINO™. | |n243-img1| | +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ + | `244-named-entity-recognition `__ |br| |c244| | Named entity recognition with OpenVINO™. | | + +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ + | `245-typo-detector `__ | English Typo Detection in sentences with OpenVINO™. | |n245-img1| | + +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ + | `246-depth-estimation-videpth `__ | Monocular Visual-Inertial Depth Estimation with OpenVINO™. | |n246-img1| | + +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ + | `247-code-language-id `__ |br| |n247| | Identify the programming language used in an arbitrary code snippet. | | + +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ + | `248-stable-diffusion-xl `__ | Image generation with Stable Diffusion XL and OpenVINO™. | |n248-img1| | + +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ + | `249-oneformer-segmentation `__ | Universal segmentation with OneFormer and OpenVINO™. | |n249-img1| | + +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ -.. raw:: html - - - -`Model Training`_ ------------------- +Model Training +-------------------- Tutorials that include code to train neural networks. @@ -272,19 +270,13 @@ Tutorials that include code to train neural networks. +======================================================================================================================================+============================================================================================================================================+===========================================+ | `301-tensorflow-training-openvino `__ | Train a flower classification model from TensorFlow, then convert to OpenVINO IR. | |n301-img1| | +--------------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ -| `301-tensorflow-training-openvino-nncf `__ | Use Post-training Optimization Tool (POT) to quantize the flowers model. | | -+--------------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ | `302-pytorch-quantization-aware-training `__ | Use Neural Network Compression Framework (NNCF) to quantize PyTorch model. | | +--------------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ | `305-tensorflow-quantization-aware-training `__ |br| |c305| | Use Neural Network Compression Framework (NNCF) to quantize TensorFlow model. | | +--------------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ -.. raw:: html - - - -`Live Demos`_ ---------------- +Live Demos +-------------------- Live inference demos that run on a webcam or video files. @@ -308,12 +300,8 @@ Live inference demos that run on a webcam or video files. +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ -.. raw:: html - - - -`Recommended Tutorials`_ --------------------------- +Recommended Tutorials +--------------------- The following tutorials are guaranteed to provide a great experience with inference in OpenVINO: @@ -342,34 +330,25 @@ The following tutorials are guaranteed to provide a great experience with infere +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ | `Dolly v2 `__ | Instruction following using Databricks Dolly 2.0 and OpenVINO™. | |n240-img1| | +-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ +| `248-stable-diffusion-xl `__ | Image generation with Stable Diffusion XL and OpenVINO™. | |n248-img1| | ++-------------------------------------------------------------------------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------+ -------------------- - .. note:: If there are any issues while running the notebooks, refer to the **Troubleshooting** and **FAQ** sections in the :doc:`Installation Guide ` or start a GitHub `discussion `__. - -.. raw:: html - - - -`Additional Resources`_ -------------------------- +Additional Resources +-------------------- * `OpenVINO™ Notebooks - Github Repository `_ * `Binder documentation `_ * `Google Colab `__ -.. raw:: html - - - -`Contributors`_ --------------------------- +Contributors +-------------------- |contributors| @@ -452,6 +431,8 @@ Made with `contributors-img `__. :target: https://user-images.githubusercontent.com/41332813/158430181-05d07f42-cdb8-4b7a-b7dc-e7f7d9391877.png .. |n218-img1| image:: https://user-images.githubusercontent.com/47499836/163544861-fa2ad64b-77df-4c16-b065-79183e8ed964.png :target: https://user-images.githubusercontent.com/47499836/163544861-fa2ad64b-77df-4c16-b065-79183e8ed964.png +.. |n220-img1| image:: https://user-images.githubusercontent.com/51917466/254583163-3bb85143-627b-4f02-b628-7bef37823520.png + :target: https://user-images.githubusercontent.com/51917466/254583163-3bb85143-627b-4f02-b628-7bef37823520.png .. |n222-img1| image:: https://user-images.githubusercontent.com/18904157/166343139-c6568e50-b856-4066-baef-5cdbd4e8bc18.png :target: https://user-images.githubusercontent.com/18904157/166343139-c6568e50-b856-4066-baef-5cdbd4e8bc18.png .. |n223-img1| image:: https://user-images.githubusercontent.com/91228207/185105225-0f996b0b-0a3b-4486-872d-364ac6fab68b.png @@ -464,7 +445,9 @@ Made with `contributors-img `__. :target: https://raw.githubusercontent.com/WongKinYiu/yolov7/main/figure/horses_prediction.jpg .. |n227-img1| image:: https://user-images.githubusercontent.com/29454499/204548693-1304ef33-c790-490d-8a8b-d5766acb6254.png :target: https://user-images.githubusercontent.com/29454499/204548693-1304ef33-c790-490d-8a8b-d5766acb6254.png -.. |n228-img1| image:: https://user-images.githubusercontent.com/29454499/207795060-437b42f9-e801-4332-a91f-cc26471e5ba2.png +.. |n228-img1| image:: https://camo.githubusercontent.com/8beb0eedc6a3bcafc397399d55a7e7da4184c1c799e6351a07a7c4aef534ffc4/68747470733a2f2f757365722d696d616765732e67697468756275736572636f6e74656e742e636f6d2f32393435343439392f3230373737333438312d64373763616366382d366364632d343736352d613331622d6131363639343736643632302e706e67 + :target: https://camo.githubusercontent.com/8beb0eedc6a3bcafc397399d55a7e7da4184c1c799e6351a07a7c4aef534ffc4/68747470733a2f2f757365722d696d616765732e67697468756275736572636f6e74656e742e636f6d2f32393435343439392f3230373737333438312d64373763616366382d366364632d343736352d613331622d6131363639343736643632302e706e67 +.. |n228-img2| image:: https://user-images.githubusercontent.com/29454499/207795060-437b42f9-e801-4332-a91f-cc26471e5ba2.png :target: https://user-images.githubusercontent.com/29454499/207795060-437b42f9-e801-4332-a91f-cc26471e5ba2.png .. |n229-img1| image:: https://user-images.githubusercontent.com/95271966/206130638-d9847414-357a-4c79-9ca7-76f4ae5a6d7f.png :target: https://user-images.githubusercontent.com/95271966/206130638-d9847414-357a-4c79-9ca7-76f4ae5a6d7f.png @@ -496,6 +479,14 @@ Made with `contributors-img `__. :target: https://user-images.githubusercontent.com/29454499/244291912-bbc6e08c-c0a9-41fe-bc2d-5f89a0d2463b.png .. |n243-img1| image:: https://user-images.githubusercontent.com/29454499/251085926-14045ebc-273b-4ccb-b04f-82a3f7811b87.gif :target: https://user-images.githubusercontent.com/29454499/251085926-14045ebc-273b-4ccb-b04f-82a3f7811b87.gif +.. |n245-img1| image:: https://user-images.githubusercontent.com/80534358/224564463-ee686386-f846-4b2b-91af-7163586014b7.png + :target: https://user-images.githubusercontent.com/80534358/224564463-ee686386-f846-4b2b-91af-7163586014b7.png +.. |n246-img1| image:: https://raw.githubusercontent.com/alexklwong/void-dataset/master/figures/void_samples.png + :target: https://raw.githubusercontent.com/alexklwong/void-dataset/master/figures/void_samples.png +.. |n248-img1| image:: https://user-images.githubusercontent.com/29454499/258651862-28b63016-c5ff-4263-9da8-73ca31100165.jpeg + :target: https://user-images.githubusercontent.com/29454499/258651862-28b63016-c5ff-4263-9da8-73ca31100165.jpeg +.. |n249-img1| image:: https://camo.githubusercontent.com/f46c3642d3266e9d56d8ea8a943e67825597de3ff51698703ea2ddcb1086e541/68747470733a2f2f6769746875622d70726f64756374696f6e2d757365722d61737365742d3632313064662e73332e616d617a6f6e6177732e636f6d2f37363136313235362f3235383634303731332d66383031626430392d653932372d346162642d616132662d3939393064653463616638642e676966 + :target: https://camo.githubusercontent.com/f46c3642d3266e9d56d8ea8a943e67825597de3ff51698703ea2ddcb1086e541/68747470733a2f2f6769746875622d70726f64756374696f6e2d757365722d61737365742d3632313064662e73332e616d617a6f6e6177732e636f6d2f37363136313235362f3235383634303731332d66383031626430392d653932372d346162642d616132662d3939393064653463616638642e676966 .. |n301-img1| image:: https://user-images.githubusercontent.com/15709723/127779607-8fa34947-1c35-4260-8d04-981c41a2a2cc.png :target: https://user-images.githubusercontent.com/15709723/127779607-8fa34947-1c35-4260-8d04-981c41a2a2cc.png .. |n401-img1| image:: https://user-images.githubusercontent.com/4547501/141471665-82b28c86-cf64-4bfe-98b3-c314658f2d96.gif @@ -538,7 +529,7 @@ Made with `contributors-img `__. :target: https://mybinder.org/v2/gh/openvinotoolkit/openvino_notebooks/HEAD?filepath=notebooks%2F101-tensorflow-to-openvino%2F101-tensorflow-to-openvino.ipynb .. |c102| image:: https://camo.githubusercontent.com/84f0493939e0c4de4e6dbe113251b4bfb5353e57134ffd9fcab6b8714514d4d1/68747470733a2f2f636f6c61622e72657365617263682e676f6f676c652e636f6d2f6173736574732f636f6c61622d62616467652e737667 :width: 109 - :target: https://colab.research.google.com/github/eaidova/openvino_notebooks/blob/ea/pt_tutorial/notebooks/102-pytorch-to-openvino/102-pytorch-to-openvino.ipynb + :target: https://colab.research.google.com/github/openvinotoolkit/openvino_notebooks/blob/main/notebooks/102-pytorch-to-openvino/102-pytorch-to-openvino.ipynb .. |n103| image:: https://mybinder.org/badge_logo.svg :target: https://mybinder.org/v2/gh/openvinotoolkit/openvino_notebooks/HEAD?filepath=notebooks%2F103-paddle-onnx-to-openvino-classification%2F103-paddle-onnx-to-openvino-classification.ipynb .. |n104| image:: https://mybinder.org/badge_logo.svg @@ -571,6 +562,11 @@ Made with `contributors-img `__. .. |c120| image:: https://camo.githubusercontent.com/84f0493939e0c4de4e6dbe113251b4bfb5353e57134ffd9fcab6b8714514d4d1/68747470733a2f2f636f6c61622e72657365617263682e676f6f676c652e636f6d2f6173736574732f636f6c61622d62616467652e737667 :width: 109 :target: https://colab.research.google.com/github/openvinotoolkit/openvino_notebooks/blob/main/notebooks/120-tensorflow-object-detection-to-openvino/120-tensorflow-object-detection-to-openvino.ipynb +.. |n121| image:: https://mybinder.org/badge_logo.svg + :target: https://mybinder.org/v2/gh/openvinotoolkit/openvino_notebooks/HEAD?filepath=notebooks%2F121-convert-to-openvino%2F121-convert-to-openvino.ipynb +.. |c121| image:: https://camo.githubusercontent.com/84f0493939e0c4de4e6dbe113251b4bfb5353e57134ffd9fcab6b8714514d4d1/68747470733a2f2f636f6c61622e72657365617263682e676f6f676c652e636f6d2f6173736574732f636f6c61622d62616467652e737667 + :width: 109 + :target: https://colab.research.google.com/github/openvinotoolkit/openvino_notebooks/blob/main/notebooks/121-convert-to-openvino/121-convert-to-openvino.ipynb .. |n209| image:: https://mybinder.org/badge_logo.svg :target: https://mybinder.org/v2/gh/openvinotoolkit/openvino_notebooks/HEAD?filepath=notebooks%2F209-handwritten-ocr%2F209-handwritten-ocr.ipynb .. |n201| image:: https://mybinder.org/badge_logo.svg @@ -617,6 +613,11 @@ Made with `contributors-img `__. :target: https://mybinder.org/v2/gh/openvinotoolkit/openvino_notebooks/HEAD?labpath=notebooks%2F218-vehicle-detection-and-recognition%2F218-vehicle-detection-and-recognition.ipynb .. |n219| image:: https://mybinder.org/badge_logo.svg :target: https://mybinder.org/v2/gh/openvinotoolkit/openvino_notebooks/HEAD?labpath=notebooks%2F219-knowledge-graphs-conve%2F219-knowledge-graphs-conve.ipynb +.. |n220| image:: https://mybinder.org/badge_logo.svg + :target: https://mybinder.org/v2/gh/openvinotoolkit/openvino_notebooks/HEAD?labpath=notebooks%2F220-cross-lingual-books-alignment%2F220-cross-lingual-books-alignment.ipynb +.. |c220| image:: https://camo.githubusercontent.com/84f0493939e0c4de4e6dbe113251b4bfb5353e57134ffd9fcab6b8714514d4d1/68747470733a2f2f636f6c61622e72657365617263682e676f6f676c652e636f6d2f6173736574732f636f6c61622d62616467652e737667 + :width: 109 + :target: https://colab.research.google.com/github/openvinotoolkit/openvino_notebooks/blob/main/notebooks/220-cross-lingual-books-alignment/220-cross-lingual-books-alignment.ipynb .. |n221| image:: https://mybinder.org/badge_logo.svg :target: https://mybinder.org/v2/gh/openvinotoolkit/openvino_notebooks/HEAD?labpath=notebooks%2F221-machine-translation%2F221-machine-translation.ipynb .. |c221| image:: https://camo.githubusercontent.com/84f0493939e0c4de4e6dbe113251b4bfb5353e57134ffd9fcab6b8714514d4d1/68747470733a2f2f636f6c61622e72657365617263682e676f6f676c652e636f6d2f6173736574732f636f6c61622d62616467652e737667 @@ -642,6 +643,10 @@ Made with `contributors-img `__. :target: https://mybinder.org/v2/gh/openvinotoolkit/openvino_notebooks/HEAD?filepath=notebooks%2F243-tflite-selfie-segmentation%2F243-tflite-selfie-segmentation.ipynb .. |c243| image:: https://camo.githubusercontent.com/84f0493939e0c4de4e6dbe113251b4bfb5353e57134ffd9fcab6b8714514d4d1/68747470733a2f2f636f6c61622e72657365617263682e676f6f676c652e636f6d2f6173736574732f636f6c61622d62616467652e737667 :target: https://colab.research.google.com/github/openvinotoolkit/openvino_notebooks/blob/main/notebooks/243-tflite-selfie-segmentation/243-tflite-selfie-segmentation.ipynb +.. |c244| image:: https://camo.githubusercontent.com/84f0493939e0c4de4e6dbe113251b4bfb5353e57134ffd9fcab6b8714514d4d1/68747470733a2f2f636f6c61622e72657365617263682e676f6f676c652e636f6d2f6173736574732f636f6c61622d62616467652e737667 + :target: https://colab.research.google.com/github/openvinotoolkit/openvino_notebooks/blob/main/notebooks/244-named-entity-recognition/244-named-entity-recognition.ipynb +.. |n247| image:: https://mybinder.org/badge_logo.svg + :target: https://mybinder.org/v2/gh/openvinotoolkit/openvino_notebooks/HEAD?filepath=notebooks%2F247-code-language-id%2F247-code-language-id.ipynb .. |c305| image:: https://camo.githubusercontent.com/84f0493939e0c4de4e6dbe113251b4bfb5353e57134ffd9fcab6b8714514d4d1/68747470733a2f2f636f6c61622e72657365617263682e676f6f676c652e636f6d2f6173736574732f636f6c61622d62616467652e737667 :width: 109 :target: https://colab.research.google.com/github/openvinotoolkit/openvino_notebooks/blob/main/notebooks/305-tensorflow-quantization-aware-training/305-tensorflow-quantization-aware-training.ipynb