* choose Pcore to compile model for GPU plugin
* provide function to update executor config
* set callback executor to nullptr for GPU plugin
* fix code style
* fix warning
* optimize duplicate code
* set callback executor to nullptr for another gpu compile_model
* add description for new function
* add smoke test
* fix code style
* modify function definition
---------
Co-authored-by: Wanglei Shen <wanglei.shen@intel.com>
* Add static shape adapter
- Adapters holds CPU dimension which can be reference to it or vector
- Add ov::optional for holding optional result from shape inference
- Add new `infer` function in `IStaticShapeInfer`
* Temporary support of StaticShape
* Minor corrections in ShapeInferenceTA
* Migrate shape_infer to new interface version
* Replace StaticShape by adapter implementation
* Replace IShapeInferCommon by IStaticShapeInfer
* Correct code formatting
* Fix build issues
* NodeValidationFailure::create for StaticShapeRef
* Review shape inference for reshape operator
- review shape_infer implementation
- add more unit test for static and dynamic shapes
* Fix build issues
* Correct minus one dim calculation
* Fix build issues on windows
* Improve resolving special minus one
* Use NODE_SHAPE_INFER_CHECK
* Update product in/out calculations
* Temporary add ngraph header to solve build issue
* Correct minus one dim calc when static part same
* Add check for scalar input
* Remove debug message
* Fix `minus one` dynamic dimension calculation
* Fix `minus one` dynamic dimension calculation
* Fix merge issues in reshape
Minor refactor reshape evaluate
* Don't pass input label on minus one pattern
when input dimension will be modified.
* [GPU] Fix outputs are not allocated in loop_inst
* Fill empty padding when the number of output paddings is less than num_outputs
* Fill empty data types when the number of output data types is less than num_outputs
* Modify postprocess_output_memory to set output memory without set_output_memory function
* In postprocess_output_memory, get concatenated_output_mem using input_info including output idx
* Modify gpu functional tests for dynamic loop to check multiple outputs of dynamic loop
* update postprocessing for condition
* Fix empty dimension issue for scalar value
* change code to get output paddings and output data type in primitive
* allocate memory for scalar data type with zero dimension
* Fix mismatch issue of input layout with shape and data types in body_network
* Fix output setting in post-processing
* pass bytes_count to gpu_usm params
* Fix condition gpu functional test issue
* Revert "allocate memory for scalar data type with zero dimension"
This reverts commit 2f10f3687c78406b20d52b6e37b1be2a30b4b73f.
* reinterpret one dimension memory buffer to zer dimension memor buffer to avoid zero byte memory allocation issue
* skip excessive mem alloc request in build
* update mem check function
* fix os behavior
* update mem size check location
* only dynamic shape case takes check_allocatable
* update check condition
* Add Rotation support to primitive and kernel
* Add unit tests
* Add transformation for NMSRotated
* add single-layer tests
* Fix: angle value for the same box may have its sign changed several times passing through iterations of batch and class loops.
* fix review comments
Transformation fuses Transpose on first or second MatMul's input
and sets MatMul's transpose_a/transpose_b accordingly.
TransposeMatMul is already part of SmartReshape, but it can be added
to MOCTransformations as well so native models that are don't use reshape
can benefit from that.
Ticket: CVS-118908
* Initial implementation of primitive, kernel selector, dummy kernel for RMS Norm
Signed-off-by: Andrew Park <andrew.park@intel.com>
* RMS ref kernel implementation with single WI
Signed-off-by: Andrew Park <andrew.park@intel.com>
* Add TC and reference func for ov_gpu_unit_tests
Signed-off-by: Andrew Park <andrew.park@intel.com>
* Add internal RMS norm op
Signed-off-by: Andrew Park <andrew.park@intel.com>
* Add transformation which fuse RMS decompsition pattern to RMS internal op
Signed-off-by: Andrew Park <andrew.park@intel.com>
* Fix pattern for RMS fusion transformation
* Update rms ref kernel for optimization and additional planar format suuport
* Initial impl for optimized rms kernel excluding leftovers handling and case smaller than vector size
* Update the initial version to handle leftovers and case smaller than vector size
* Fuse pre decom and post comp reorders additionally
* Enable dynamic impl for rms again
* Revert fuse pre decomp and post comp reorders additionally
* Add subgraph TC for ov_gpu_func_tests
* decrease error margin for f32 data type
* update description
Signed-off-by: Andrew Park <andrew.park@intel.com>
* update test param for input shapes
* Apply comments
* Fix failed TC for invalid gamma element type
* Apply comments
Signed-off-by: Andrew Park <andrew.park@intel.com>
* Update pattern that fuse post reorder together
* Apply comments
---------
Signed-off-by: Andrew Park <andrew.park@intel.com>
* Gather needs to keep the original input/output rank
- because the parameters as indices, batch_dims and axis depend on the rank.
- add input_rank to gather primitive.
* don't query on set_preferred_formats pass
-when the force_implementations is set.
-when forcing_impl is not onednn.
* [GPU] Fixed data generation for f16 fusion tests
* [GPU] Temporary tolerance increase for failed tests on iGPU
* [GPU] Temporary skip or tolerance increase for failed tests on dGPU
* Add group_normalization_kernel_selector
* Define group_normalization GPU primitive and its instantiation
* Add GroupNormalization operation builder
* Add test class for GroupNormalization operator
* Add instantiation of GroupNormalization test for GPU Plugin
* Disable GroupNormalizationDecomposition transformation in GPU Plugin
* Add GroupNormalizationKernelRef implementation
* Add GroupNormalization unit tests which cover blocked layout support
* [GPU] enable dynamic loop
- support multiple outputs
- support dynamic loop memory allocation
- support negative num_iterations
- implement calc_output_layouts
- add dynamic loop functional / unit tests
* Fix fail to check memory to set when original 1d data
- follow up code reviews
* Fix unit test failures
* Follow up code review
* Modify concat memory map creation process
* Check whether or not first input of loop is num_iteration_id
* Follow up code review
- refactoring preprocess_backedge_memory
* * Fix ci failures
* Clear custom_outputs_vec for condition
* Add num_outputs for condition and loop
* *Fix constant and param of body network have mismatched layouts
* Set consts.needsBatchInterpretation for const
* * refactoring is_dynamic in loop_inst::execute
* * remove wait_for_events in body_network execution loop
* * Remove redundant events
* * follow-up code review - modify OPENVNO_ASSERT
* * Remove redundant codes in loop_inst::execute
* * add current iteration update nodes into the ov::Model
* * rollback some codes for the performance degradation