Maruf Ahmed: Evaluation of NVIDIA H200 GPUs for training Large-scale AI/ML models at NCI

Maruf Ahmed: Evaluation of NVIDIA H200 GPUs for training Large-scale AI/ML models at NCI

In recent years, Machine Learning (ML) has been used alongside traditional physics and numerical-method-based models. The revolution in Artificial Intelligence (AI) has impacted many areas of research, and new AI models are being released regularly. The AI models are being used in major scientific domains, like atmospheric modelling, weather prediction, ocean wave simulation, protein structure folding, and drug discovery process, to name a few. Recently, the National Computational Infrastructure (NCI) has acquired a new generation of Hopper H200 GPUs, which are equipped with advanced Tensor cores and larger available memory, enabling higher throughput training and inference workloads. The H200 has almost 4 times more memory than the V100. Furthermore, the GPUs are connected with faster NV-Links for fast communication. The State-of-the-Art (SOTA) ML models require a huge amount of time and computing resources; thus, it is infeasible to train such models without modern GPUs. With the introduction of H200 GPUs, NCI is now capable of training large-scale SOTA AI/ML models. To showcase the capabilities of H200, we experimented with two SOTA ML models for training, which are the ECMWF Artificial Intelligence Forecasting System (AIFS) and Nvidia FourCastNet 3. Our experiments show that the H200 is perfectly capable of training SOTA AI models.