MGARD-GPU is a CUDA implementation of the MGARD lossy compressor, which significantly improves MGARD's throughput on GPU-based heterogeneous HPC systems.
- Double and single precision data types
- High dimensional data (upto 4D)
- Compression with S-norm
- Uniform and non-uniform data
- End-to-end high performance compression on GPUs
- Performance pre-tuned for Volta and Turing GPUs
- NVIDIA GPUs ( tested on Volta, Turing)
- CUDA 11.0+
- GCC 7.4.0+
- CMake 3.19+
- Option 1: one-step configure and build MGARD with NVCOMP:
build_mgard_cuda.sh - Option 2: manual confiugre and build
-
Step 1: configure and build NVCOMP.
-
Step 2: configure MGARD as follows:
cmake -S <MGARD_SRC_DIR> -B <MGARD_BUILD_DIR> -DMGARD_ENABLE_CUDA=ON -DNVCOMP_ROOT=<NVCOMP_INSTALL_DIR> -
Step 3: build MGARD:
cmake --build <MGARD_BUILD_DIR> -j8
-
-
Step 1: Include the header file. MGARD-GPU APIs are included in both
mgard/compress.hppandmgard/mgard_cuda_api.h.- Use
mgard/compress.hppif the user programs are to be compiled with C/C++ compilers. - Use
mgard/mgard_cuda_api.hif the user programs are to be compiled with CUDA compilers.
- Use
-
Step 2: Initialize mgard_cuda::Handle. An object
mgard_cuda::Handleneeds to be created and initialized. This initializes the necessary environment for efficient compression on the GPU. It only needs to be created once if the input shape is not changed. For example, compressing the same variable on different timesteps only needs the handle to be created once. Also, the same handle can be shared in between compression and decompression APIs.- For uniform grids:
mgard_cuda::Handle<N_dims, D_type>(std::vector<size_t> shape).[In] D_type: Input data type (float or double).[In] N_dims: Total number of dimensions (<=4)[In] shape: Stores the size in each dimension (from slowest to fastest).
- For non-uniform grids:
mgard_cuda::Handle<N_dims, D_type>(std::vector<size_t> shape, std::vector<T*> coords).[In] coords: The coordinates in each dimension (from slowest to fastest).
- For uniform grids:
-
Step 3: Use mgard_cuda::Array.
mgard_cuda::Arrayis used for holding a managed array on GPU.- For creating an array.
mgard_cuda::Array::Array<N_dims, D_type>(std::vector<size_t> shape)creates an manged array on GPU withshape. - For loading data into an array.
void mgard_cuda::Array::loadData(D_type *data, size_t ld = 0)copiesdatainto the the managed array on GPU.datacan be on either on CPU or GPU. An optionalldcan be provided for specifying the size of the leading dimension. - For accessing data from CPU
D_type * mgard_cuda::Array::getDataHost()returns a CPU pointer of the array. - For accessing data from GPU
D_type * mgard_cuda::Array::getDataDevice(size_t &ld)returns a GPU pointer of the array with the leading dimension. - For getting the shape of an array.
std::vector<size_t> mgard_cuda::Array::getShape()returns the shape of the managed array.
Note:
mgard_cuda::Arraywill automatically release its internal CPU/GPU array when it goes out of scope. - For creating an array.
-
Step 4: Call compression/decompression API.:
- For compression:
mgard_cuda::Array<1, unsigned char> mgard_cuda::compress(mgard_cuda::Handle<N_dims, D_type> &handle, mgard_cuda::Array<N_dims, D_type> in_array, mgard_cuda::error_bound_type type, D_type tol, D_type s)[In] in_array: Input data to be compressed (its value will be altered during compression).[In] type: Error bound type.mgard_cuda::RELfor relative error bound ormgard_cuda::ABSfor absolute error bound.[In] tol: Error bound.[In] s: Smoothness parameter.[Return]: Compressed data.
- For decompression:
mgard_cuda::Array<N_dims, D_type> mgard_cuda::decompress(mgard_cuda::Handle<N_dims, D_type> &handle, mgard_cuda::Array<1, unsigned char> compressed_data)[In] compressed_data: Compressed data.[Return]: Decompressed data.
- For compression:
-
Optimize for specific GPU architectures: MGARD-GPU is pre-tuned for Volta and Turing GPUs. To enable this optimization, the following additional CMake options need to be enabled when configuring MGARD-GPU.
- For Volta GPUs:
-DMGARD_ENABLE_CUDA_OPTIMIZE_VOLTA=ON - For Turing GPUs:
-DMGARD_ENABLE_CUDA_OPTIMIZE_TURING=ON
- For Volta GPUs:
-
Optimize for GPUs on edge systems: MGARD-GPU capable of using FMA instructions that can help improve the compression/decompression performance on consumer-class GPUs (GeForce GPUs on edge systems), where they have relative low throughput on some arithmetic operations. To enable this optimization, enable the following option when configuring MGARD-GPU.
-MGARD_ENABLE_CUDA_FMA=ON
-
Optimize for fast CPU-GPU data transfer: It is recommanded to use pinned memory on CPU for loading data into
mgard_cuda::Arraysuch that it can enable fast CPU-GPU data transfer.- To allocate pinned memory on CPU:
mgard_cuda::cudaMallocHostHelper(void ** data_ptr, size_t size). - To free pinned memory on CPU:
mgard_cuda::cudaFreeHostHelper(void * data_ptr)
- To allocate pinned memory on CPU:
-
Selecting the target GPU: By default MGARD-GPU uses the first GPU (i.e., DEVICE_ID=0) in a multi-GPU system as the target GPU for compression/decompression. To configure MGARD-GPU to use a differet GPU.
mgard_cuda::Configcan be used to set the target GPU and pass its instance tomgard_cuda::Handle. Then, all compression/decompression routines will use the target GPU that corresponds to the one set inmgard_cuda::Handle.mgard_cuda::Config config; config.dev_id = <GPU ID>; mgard_cuda::Handle<3, double>(shape, config);
The following code shows how to compress/decompress a 3D dataset.
#include <vector>
#include <iostream>
#include "mgard/compress.hpp"
int main()
{
size_t n1 = 10;
size_t n2 = 20;
size_t n3 = 30;
//prepare
std::cout << "Preparing data...";
double * in_array_cpu;
mgard_cuda::cudaMallocHostHelper((void **)&in_array_cpu, sizeof(double)*n1*n2*n3);
//... load data into in_array_cpu
std::vector<size_t> shape{ n1, n2, n3 };
mgard_cuda::Handle<3, double> handle(shape);
mgard_cuda::Array<3, double> in_array(shape);
in_array.loadData(in_array_cpu);
std::cout << "Done\n";
std::cout << "Compressing with MGARD-GPU...";
double tol = 0.01, s = 0;
mgard_cuda::Array<1, unsigned char> compressed_array = mgard_cuda::compress(handle, in_array, mgard_cuda::REL, tol, s);
size_t compressed_size = compressed_array.getShape()[0]; //compressed size in number of bytes.
unsigned char * compressed_array_cpu = compressed_array.getDataHost();
std::cout << "Done\n";
std::cout << "Decompressing with MGARD-GPU...";
// decompression
mgard_cuda::Array<3, double> decompressed_array = mgard_cuda::decompress(handle, compressed_array);
mgard_cuda::cudaFreeHostHelper(in_array_cpu);
double * decompressed_array_cpu = decompressed_array.getDataHost();
std::cout << "Done\n";
}
- A comprehensive example about how to use MGARD-GPU is located in here.