Commit Graph

557 Commits

Author SHA1 Message Date
i-robot 2e8eac8341
!30367 auto_parallel_adasum_support_data_parallel
Merge pull request !30367 from yao_yf/auto_parallel_adasum_support_data_parallel
2022-02-23 06:50:26 +00:00
jiahongQian 8a2151d8bb allgather reducescatter fusion 2022-02-23 10:47:30 +08:00
i-robot 5bee7156b9
!30369 add_virtualdataset_ut
Merge pull request !30369 from lilei/add_virtualdataset_ut
2022-02-23 02:21:34 +00:00
yao_yf 19236b1a70 auto parallel adasum support data parallel and hybrid parallel 2022-02-22 20:14:00 +08:00
huangxinjing 092ba035e3 Add global norm parallel support 2022-02-22 16:08:53 +08:00
yanghaoran bfe139b662 takedonw test_auto_parallel_adasum.py to ensure gate stability 2022-02-22 14:51:03 +08:00
lilei bc62e24d94 add_virtualdataset_ut 2022-02-22 11:38:10 +08:00
i-robot 94c8c6355c
!30294 auto_parallel_adasum_checks_and_ut.
Merge pull request !30294 from yao_yf/auto_parallel_adasum_checks_and_ut
2022-02-22 01:10:24 +00:00
huangxinjing 5e325ac336 [AUTO_PARALLEL]Fix insert nodes error 2022-02-21 16:39:50 +08:00
i-robot 7386612515
!29820 moe_topk routing
Merge pull request !29820 from wangshengnan123/moe_topk_routing
2022-02-21 06:21:55 +00:00
yao_yf 4b79d4c425 auto parallel adasum uts and checks 2022-02-21 10:41:09 +08:00
wangshengnan123 7322426648 top_k routing 2022-02-18 14:10:22 +08:00
i-robot 48d4f34576
!30167 Add UT case for 'Convolution+Transformer' structure
Merge pull request !30167 from Bert0108/ut_conformer
2022-02-17 07:43:58 +00:00
Bert0108 25a9c73a08 add ut case for monitoring the conformer structure 2022-02-17 15:08:01 +08:00
liuluobin b0b79ef8fb Add backend check for RandomChoiceWithMask 2022-02-16 18:56:57 +08:00
i-robot d5100190c9
!29612 Impementatin of CropAndResizeInfo and ROIAlignInfo.
Merge pull request !29612 from liuluobin/crop_resize_master
2022-02-15 11:36:06 +00:00
liuluobin 80fd640819 Implementation of ROIAlignInfo and CropAndResizeInfo 2022-02-14 10:46:09 +08:00
Margaret_wangrui 462c38813d [ME] Add parameter name check. 2022-02-14 10:07:43 +08:00
yangzhenzhang 4061c11e18 support pad mode for conv2d transpose 2022-02-09 11:52:47 +08:00
i-robot a55b1b5e05
!29453 support pad mode for conv2d parallel operator
Merge pull request !29453 from yangzhenzhang/support-pad-mode-for-conv2d
2022-02-07 06:58:00 +00:00
i-robot ea391f6eb1
!29505 Fix AllGather Cast when the parameters are shared
Merge pull request !29505 from huangxinjing/fx_allgather_cast
2022-01-29 07:46:10 +00:00
i-robot b0b3d7ede7
!29421 Implementation of BoundingBoxEncodeInfo, IOUInfo and RandomChoickWithMaskInfo
Merge pull request !29421 from liuluobin/parallel_ops_master
2022-01-26 06:40:32 +00:00
i-robot 55ba926a04
!29087 optimize sharding propagation&add ut on pangu
Merge pull request !29087 from bichaoyang/master
2022-01-25 11:59:25 +00:00
huangxinjing 34101d748d Add cast ut and fix juge condition 2022-01-25 18:14:11 +08:00
yangzhenzhang 6df5e9aa5d support pad mode for conv2d 2022-01-25 17:26:47 +08:00
liuluobin fb79c508f2 Implementation of BoundingBoxEncodeInfo, IOUInfo, RandomChoiceWithMaskInfo. 2022-01-25 15:49:48 +08:00
b00518648 2ff3425c76 1.consider comm cost inside a op when select a strategy;2.deal with the ops that share same param;3.add ut to fix the perfermence at pangu_alpha 2022-01-25 15:05:37 +08:00
yangzhenzhang 4d0b1a47ee support config group for batchnorm 2022-01-22 09:29:52 +08:00
yangzhenzhang 000c5b5394 support dilation for conv2d 2022-01-20 19:04:52 +08:00
i-robot 7bb5819889
!29265 fix resize_bilinear infer
Merge pull request !29265 from jiangzhenguang/resize_bilinear
2022-01-20 06:51:03 +00:00
jiangzhenguang abddc993ea fix resize_bilinear 2022-01-18 20:33:44 +08:00
yangzhenzhang 5514189257 support group for conv2d 2022-01-18 17:28:13 +08:00
yangzhenzhang 6dd7333c0b fix bug for conv2d transpose 2022-01-14 16:36:03 +08:00
yangzhenzhang a9d2e255f5 support single direction exchange for conv2d 2022-01-13 09:15:49 +08:00
i-robot 017cb5f3ad
!27980 auto insert VirtualDataset node for master
Merge pull request !27980 from lilei/insert_virtualdataset_for_master
2022-01-12 01:14:44 +00:00
yangzhenzhang e5df74e9e4 compute top bottom overlap for conv2d 2022-01-10 16:15:04 +08:00
Xiaoda Zhang 6d8320fa66 1) fix the exact division in moe;
2) changing CumSum from composition to a single Operator;
3) add InferMirrorOps for CumSumInfo.
2021-12-30 19:20:05 +08:00
Xiaoda Zhang 1bdb610b34 changing default value of single-loop flag 2021-12-28 17:56:30 +08:00
i-robot dd90a56d68 !28073 fix code warning && remove save_graphs use in st/ut
Merge pull request !28073 from huanghui/fix-warning
2021-12-24 09:29:25 +00:00
i-robot 22c25ec10e !27862 [Auto parallel] [Sharding propagation] dealing with cast
Merge pull request !27862 from Xiaoda/119-adapting-sharding-propagation
2021-12-24 07:33:47 +00:00
huanghui 74ca50e652 fix code warning && remove save_graphs use in st/ut 2021-12-24 14:45:37 +08:00
Xiaoda Zhang 66c7474e5a remove CastInfo from CNODE 2021-12-23 19:05:02 +08:00
zhuyuxiao dd7bbf92dd change API 2021-12-23 14:26:06 +08:00
lilei 017aa359a6 insert VirtualDataset node for master 2021-12-23 09:25:05 +08:00
i-robot 2fbec9a554 !27856 use neighbor-exchange-v2 for conv2d
Merge pull request !27856 from yangzhenzhang/use-neighborexchangev2-for-conv2d
2021-12-21 10:59:54 +00:00
yangzhenzhang 8a68577756 use neighbor-exchange-v2 for conv2d 2021-12-21 14:27:41 +08:00
wzw a9b78682d5 parallel ut refactor 3 2021-12-21 14:26:58 +08:00
yangzhenzhang 5f6477b022 add output strategy for gather op 2021-12-16 09:24:50 +08:00
i-robot d49f5e6caf !27525 support optimizer parallel for adafactor
Merge pull request !27525 from yangzhenzhang/support-opt-parallel-for-adafactor
2021-12-15 01:25:12 +00:00
yao_yf 30576c6a75 fix reshape bool type in auto parallel 2021-12-13 20:24:33 +08:00
yangzhenzhang 2a0b528084 support opt parallel for adafactor 2021-12-13 11:21:44 +08:00
i-robot 938dc8abd0 !27439 [Auto parallel] Add new operatorInfo for Parallel: CumSum
Merge pull request !27439 from Xiaoda/117-add-cumsum-op
2021-12-10 07:00:06 +00:00
i-robot 0e358f4cb3 !27428 revert insert VirtualDataset node for master
Merge pull request !27428 from lilei/modify_virtualdataset_for_master
2021-12-09 02:47:13 +00:00
Xiaoda Zhang 8042c88223 add the new operatorInfo for parallel: CumSum 2021-12-09 09:21:52 +08:00
lilei 2edf6ab33b revert insert VirtualDataset node for master 2021-12-08 17:39:42 +08:00
i-robot faaec746f7 !27401 add more ut tests for allreduce fusion
Merge pull request !27401 from jiahongQian/master
2021-12-08 09:38:37 +00:00
jiahongQian b03c8d18d3 add more ut tests 2021-12-08 14:17:02 +08:00
i-robot ffca7b08a5 !27237 auto insert VirtualDataset node for master
Merge pull request !27237 from lilei/modify_virtualdataset_for_master
2021-12-07 09:02:50 +00:00
i-robot f40668ef73 !27251 test_micro_batch_Interleaved
Merge pull request !27251 from lilei/add_parallel_ut
2021-12-07 02:18:09 +00:00
lilei 05189459ab auto insert VirtualDataset node for master 2021-12-06 21:39:37 +08:00
lilei e933aa268b test_micro_batch_Interleaved 2021-12-06 18:37:21 +08:00
i-robot 2d23b698a6 !27024 add allreduce fusion by size
Merge pull request !27024 from jiahongQian/master
2021-12-06 05:32:12 +00:00
q00596439 de36fdc169 add allreduce fusion size and unify the interface 2021-12-05 21:35:11 +08:00
huangxinjing 8c9b2b93a8 Add transformer 2021-12-03 09:08:57 +08:00
yangzhenzhang 7454b8f8f2 check args for shard 2021-12-02 11:51:29 +08:00
Xiaoda Zhang 364858cbc9 In sharding propagation, to keep strategy consistent of parameter being used by multiple operators, we check the edge with one node of TmpIdentityInfo 2021-11-30 18:58:33 +08:00
Xiaoda Zhang 04db51a528 In a previous PR (https://gitee.com/mindspore/mindspore/pulls/26807/), we replaced 'auto_parallel_search_mode' by 'search_mode' directly.
However, to be forward compatitable, it is suitable to keep 'auto_parallel_search_mode' available. This PR recovers the 'auto_parallel_search_mode' interface and adds a warning when using this old interface.

This PR also deals with other codestyle things.
2021-11-29 15:40:23 +08:00
i-robot 9f8ec2c5ab !26807 [Auto parallel] [Sharding propagation] Interface change of sharding propagation
Merge pull request !26807 from Xiaoda/113-auto-parallel-search-mode-changes-to-search-mode
2021-11-26 01:48:58 +00:00
i-robot 6ecbc97fd6 !26804 virtual_dataset_avoid_auto_parallel
Merge pull request !26804 from yao_yf/virtual_dataset_avoid_auto_parallel
2021-11-26 01:48:53 +00:00
i-robot b282414de7 !26619 arallel_ut_refactoring
Merge pull request !26619 from 王志伟/parallel_ut_refactoring1
2021-11-26 01:34:49 +00:00
Xiaoda Zhang ad5ac77ae8 1) 'auto_parallel_search_mode' changes to 'search_mode';
2) 'sharding_propagation' moves to 'search_mode';
2021-11-25 19:38:21 +08:00
yao_yf f29ce1fb60 virtual dataset avoid auto parallel 2021-11-25 18:47:35 +08:00
i-robot 519f14a909 !26006 slice recompute activation
Merge pull request !26006 from yao_yf/add_transformer_slice_activation_config
2021-11-24 06:14:28 +00:00
wzw 86c5ad20c8 parallel_ut_refactoring1 2021-11-24 09:41:52 +08:00
i-robot 1b8c2ff0e9 !26414 fault_recover_by_mirror_group_fix_opt_shard
Merge pull request !26414 from yao_yf/fault_recover_by_mirror_group_fix_opt_shard
2021-11-20 03:06:45 +00:00
yao_yf 188d39da83 slice_activation_in_recompute
slice recompute activation
2021-11-19 10:51:54 +08:00
yao_yf 01dc4bbdf9 fix fault recover in optimizer shard 2021-11-19 09:46:27 +08:00
Xiaoda Zhang df67e74eaf making sharding_propagation smooth, add a reshape justification:
1) when propagate sharding strategy from one op to another, try to find the strategy with zero communication cost;
2) if there is no such strategy, find the strategy with minimum communication cost, and raise a warning;
2021-11-18 15:55:50 +08:00
i-robot 9f52343a6a !26350 add check for resizenearestneighbor parallel op
Merge pull request !26350 from yangzhenzhang/add-check-for-resize-op
2021-11-18 05:45:57 +00:00
yangzhenzhang ba99e4c505 add check for resize op 2021-11-16 15:39:52 +08:00
ttudu 33ac1de062 fix bug 2021-11-16 11:14:58 +08:00
i-robot 7a73bae5c3 !26036 add output strategy for matmul operator
Merge pull request !26036 from yangzhenzhang/add-output-strategy-for-op-init
2021-11-12 01:22:22 +00:00
Xiaoda Zhang a772767265 support reshape in sharding propagation:
1) using 'swc index of strategy_cost_' as reshape's selected strategy;
2) when encountering reshape in BFS, select the 'swc index' with zero communication cost;
3) when encountering a reshape that is already visited, check whether there exists communication between reshape and current operator. It is OK if communication happens between two configured operators;
4) currently, two consecutive reshapes are not supported;
5) adjusting BFS structure in graph_costmodel.cc;
6) adjusting some code in step_auto_parallel.cc to avoid cyclomatic complexity.
2021-11-11 11:35:43 +08:00
yangzhenzhang 8431ba616c add output strategy for op init 2021-11-10 15:35:45 +08:00
huangxinjing f354ab22a3 add pipeline shard interface
Add support for no pipeline accugradient

Add delay tag for fusion op

Optimizer the visite order

add mirror for mini step control

Move the group to attributes

Add gradient_shard control for the mini step

Fix code stype

Fix ut description

Add interface
2021-11-09 09:11:13 +08:00
i-robot ded1c77bbf !25765 neighborExchangeV2 & grad
Merge pull request !25765 from TuDouNi/neighborExchangeV2
2021-11-05 09:31:03 +00:00
ttudu e953c15cd2 NeighborExchangeV2 & Grad 2021-11-05 11:00:46 +08:00
yangzhenzhang 6ad6304b77 add output strategy 2021-11-04 16:28:54 +08:00
yangzhenzhang c42081619e add parallel op for resizenearestneighbor 2021-10-28 11:31:00 +08:00
wanyiming 4fbc59a98a utfixs 2021-10-18 18:06:37 +08:00
i-robot f83070728d !24790 support user define strategy gen method under auto parallel context
Merge pull request !24790 from zhuyuxiao/master
2021-10-14 09:03:36 +00:00
zhuyuxiao 1907246931 change api 2021-10-13 10:28:30 +08:00
i-robot 3fd94000c5 !24568 Apply batch parallel in auto_parallel mode when strategies are not specified
Merge pull request !24568 from zhuyuxiao/master
2021-10-13 01:20:15 +00:00
zhuyuxiao cf76c76745 apply batch parallel in auto_parallel mode when strategies are not specified 2021-10-12 15:44:10 +08:00
yao_yf b303d6001c parallel ut refactor 2021-10-12 15:18:51 +08:00
i-robot 7cde7731b0 !23537 Update pangu reshape and softmax.
Merge pull request !23537 from linqingke/pangu
2021-09-24 08:12:13 +00:00
i-robot e7cb505e68 !23569 Produce parallel operators for ResizeBilinear and ResizeNearestNeighbor
Merge pull request !23569 from Bert0108/resizebilinear_parallel_ops
2021-09-24 03:05:23 +00:00
i-robot d37fccc56f !23544 remove deprecated gather op
Merge pull request !23544 from zhuyuxiao/master
2021-09-24 01:20:53 +00:00
Bert0108 2d3d0b673e parallel operators for ResizeBilinear and ResizeNearestNeighbor 2021-09-24 09:07:49 +08:00
linqingke acde7febef update pangu reshape and softmax performance.
Add layer norm judge

Fix layer norm name error

Fix input tyoe check

Fix ut test

Add 3d supports
2021-09-24 09:06:16 +08:00