| .. |
|
distribution
|
GPU supports p2p nccl interfaces
|
2020-11-24 14:17:53 +08:00 |
|
mpi
|
code warning clean
|
2020-09-21 09:44:39 +08:00 |
|
blocking_queue.cc
|
add push opt logic
|
2020-12-15 15:35:42 +08:00 |
|
blocking_queue.h
|
add input data type check for ps cache mode
|
2021-02-07 12:35:04 +08:00 |
|
cuda_common.h
|
add more dtypes support for gatherdgrad and other bugfix
|
2020-11-30 14:00:54 +08:00 |
|
cuda_driver.cc
|
Add device id log
|
2021-03-09 09:09:17 +08:00 |
|
cuda_driver.h
|
gpu all-reduce memory alloc fixed
|
2020-12-14 20:19:03 +08:00 |
|
cuda_env_checker.cc
|
search nvcc in entire PATH
|
2020-09-04 15:55:48 +08:00 |
|
cuda_env_checker.h
|
search nvcc in entire PATH
|
2020-09-04 15:55:48 +08:00 |
|
gpu_bucket.cc
|
add op_mul fusion based on allreduce fusion in pynative mode
|
2021-03-15 21:12:14 +08:00 |
|
gpu_bucket.h
|
add op_mul fusion based on allreduce fusion in pynative mode
|
2021-03-15 21:12:14 +08:00 |
|
gpu_buffer_mgr.cc
|
add push opt logic
|
2020-12-15 15:35:42 +08:00 |
|
gpu_buffer_mgr.h
|
add push opt logic
|
2020-12-15 15:35:42 +08:00 |
|
gpu_common.h
|
fix GPUKernelMod about the using of shard_ptr
|
2021-03-15 15:25:56 +08:00 |
|
gpu_device_address.cc
|
nlp perf(Pynative): change memory sync mode from synchronous to asynchronous in SyncHostToDevice
|
2021-02-20 02:51:51 -05:00 |
|
gpu_device_address.h
|
fix device memory leak
|
2021-01-08 16:02:54 +08:00 |
|
gpu_device_manager.cc
|
Fix GPU sync stream Segmentation fault
|
2020-12-29 16:45:49 +08:00 |
|
gpu_device_manager.h
|
Fix GPU sync stream Segmentation fault
|
2020-12-29 16:45:49 +08:00 |
|
gpu_event.cc
|
PyNative AllReduce Bucket
|
2021-03-03 15:47:18 +08:00 |
|
gpu_event.h
|
PyNative AllReduce Bucket
|
2021-03-03 15:47:18 +08:00 |
|
gpu_kernel_build.cc
|
add hardware abstract layer
|
2021-03-19 11:28:53 +08:00 |
|
gpu_kernel_build.h
|
add hardware abstract layer
|
2021-03-19 11:28:53 +08:00 |
|
gpu_kernel_runtime.cc
|
!12115 IR operators of GPU and CPU are unified as batchnorm
|
2021-03-19 14:51:26 +08:00 |
|
gpu_kernel_runtime.h
|
Support ms_function + heterogenous
|
2021-03-09 18:59:03 +08:00 |
|
gpu_launch_kernel.cc
|
add hardware abstract layer
|
2021-03-19 11:28:53 +08:00 |
|
gpu_launch_kernel.h
|
add op_mul fusion based on allreduce fusion in pynative mode
|
2021-03-15 21:12:14 +08:00 |
|
gpu_launch_mul.cc
|
add op_mul fusion based on allreduce fusion in pynative mode
|
2021-03-15 21:12:14 +08:00 |
|
gpu_launch_mul.h
|
add op_mul fusion based on allreduce fusion in pynative mode
|
2021-03-15 21:12:14 +08:00 |
|
gpu_memory_allocator.cc
|
Refactor ms_context implementation
|
2020-08-30 22:51:40 +08:00 |
|
gpu_memory_allocator.h
|
Unified code style
|
2020-07-20 10:51:55 +08:00 |
|
gpu_memory_copy_manager.cc
|
refine GPU memory swap performance
|
2020-07-21 21:37:54 +08:00 |
|
gpu_memory_copy_manager.h
|
refine GPU memory swap performance
|
2020-07-21 21:37:54 +08:00 |
|
gpu_memory_manager.cc
|
optimize the memory alloc error info
|
2021-02-04 15:30:01 +08:00 |
|
gpu_memory_manager.h
|
profiler memory
|
2021-01-19 10:21:43 +08:00 |
|
gpu_stream_assign.cc
|
fix GPUKernelMod about the using of shard_ptr
|
2021-03-15 15:25:56 +08:00 |
|
gpu_stream_assign.h
|
fix_consecutive_allreduce_bug
|
2020-08-20 16:04:56 +08:00 |
|
kernel_info_setter.cc
|
IR operators of GPU and CPU are unified as batchnorm
|
2021-03-18 19:02:28 +08:00 |
|
kernel_info_setter.h
|
IR operators of GPU and CPU are unified as batchnorm
|
2021-03-18 19:02:28 +08:00 |
|
queue_common.h
|
add trace for gpu error/excpt log
|
2020-12-04 14:03:36 +08:00 |
|
readme.md
|
mindspore path adjust
|
2020-07-14 18:07:28 +08:00 |