Why Is Setting Up a GPU Server for AI Models Critical?
Running deep learning models on a CPU is not only slow, but it becomes practically impossible for modern models with billions of parameters. A well-equipped GPU server can reduce the training time of a model from several weeks to a few hours. However, purchasing or renting hardware is only part of the story; correctly installing the driver, CUDA Toolkit, and related libraries determines your success. In this article, we will follow the complete path of setting up a GPU server from the moment you have access to the system to running your first workload, with technical details and real-world examples.
Prerequisites and Hardware Check
Before taking any action, you must ensure that your GPU is recognized by the operating system and has the appropriate driver. If you are using a cloud server (for example, ServerNet cloud services that offer GPU server rentals), you will typically receive a pre-configured image. But if you have a dedicated server, follow the steps below.
Identifying the GPU in Linux
Most GPU servers run Ubuntu 20.04 or 22.04 LTS. First, check whether the GPU is detected by the system using the following command:
lspci | grep -i nvidia
The output should look something like this:
01:00.0 3D controller: NVIDIA Corporation GA102 [GeForce RTX 3080] (rev a1)
If the output is empty, the GPU is likely not properly seated in the PCIe slot or is disabled in the BIOS. On cloud servers, this step is usually already done.
Checking the Current Driver
To see the installed driver version (if any), use the following command:
nvidia-smi
If this command does not work, the driver is not installed or not properly configured. A successful output should include the GPU name, driver version, and CUDA version (which usually comes with the driver).
Common Mistake: Many users think installing the CUDA Toolkit alone is sufficient. But the CUDA Toolkit does not work without the appropriate driver. The driver and CUDA are two separate layers that must be compatible with each other.
Installing the NVIDIA Driver
There are two main methods for installing the driver: using the official NVIDIA repository or the Ubuntu repository. The first method is recommended because it provides a newer version.
Method 1: Installing from the Official NVIDIA Repository
First, add the official repository:
sudo add-apt-repository ppa:graphics-drivers/ppa
sudo apt update
Then install the appropriate driver version. To find the latest stable version, you can search:
apt search nvidia-driver | grep -E '^nvidia-driver-[0-9]+'
The output usually includes versions like nvidia-driver-535 or nvidia-driver-545. Install version 535 or higher:
sudo apt install nvidia-driver-535
After installation, reboot the system:
sudo reboot
Method 2: Installing from the Ubuntu Repository
If the above method fails for any reason, you can use the default repository:
sudo apt install nvidia-driver-535-server
This version is optimized for servers and has greater stability.
Verifying the Driver Installation
After rebooting, run the nvidia-smi command again. The output should include a table with GPU specifications and the driver version. If you receive the error NVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver, the driver has not been loaded correctly. In this case, check the logs:
dmesg | grep -i nvidia
Troubleshooting Tip: If Secure Boot is enabled, the driver may be unsigned and fail to load. Disable Secure Boot in the BIOS or sign the driver with the MOK key.
Installing the CUDA Toolkit
The CUDA Toolkit includes the nvcc compiler, runtime libraries, and development tools. The CUDA version must be compatible with the driver version. For example, driver 535 supports CUDA 12.2.
Downloading and Installing CUDA
Visit the official CUDA Toolkit page and select the appropriate version for your operating system. For Ubuntu 22.04, run the following commands:
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt update
sudo apt install cuda-toolkit-12-2
After installation, add the CUDA paths to the environment variables. Do this by editing the ~/.bashrc file:
echo 'export PATH=/usr/local/cuda-12.2/bin:$PATH' >> ~/.bashrc
echo 'export LD_LIBRARY_PATH=/usr/local/cuda-12.2/lib64:$LD_LIBRARY_PATH' >> ~/.bashrc
source ~/.bashrc
Verifying the CUDA Installation
Check the CUDA version:
nvcc --version
The output should include a version like Cuda compilation tools, release 12.2, V12.2.140.
Installing cuDNN for Optimization
cuDNN (CUDA Deep Neural Network library) is an accelerator for neural network operations. Without it, frameworks like TensorFlow and PyTorch use slower implementations.
Installing cuDNN
To install cuDNN, you need to register on the NVIDIA website and download the appropriate files. After downloading, follow these steps:
tar -xzvf cudnn-linux-x86_64-8.9.7.29_cuda12-archive.tar.xz
sudo cp cudnn-linux-x86_64-8.9.7.29_cuda12-archive/include/cudnn*.h /usr/local/cuda-12.2/include/
sudo cp -P cudnn-linux-x86_64-8.9.7.29_cuda12-archive/lib/libcudnn* /usr/local/cuda-12.2/lib64/
sudo chmod a+r /usr/local/cuda-12.2/include/cudnn*.h /usr/local/cuda-12.2/lib64/libcudnn*
To verify the installation, check the cuDNN version:
cat /usr/local/cuda-12.2/include/cudnn_version.h | grep CUDNN_MAJOR -A 2
Installing Deep Learning Frameworks
Now that the driver, CUDA, and cuDNN are ready, it's time to install the framework. PyTorch and TensorFlow are the two main options.
Installing PyTorch with CUDA Support
To install PyTorch with CUDA 12.1 (which is compatible with CUDA 12.2), use the following command:
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
Verify that PyTorch can see the GPU:
python -c "import torch; print(torch.cuda.is_available()); print(torch.cuda.get_device_name(0))"
The output should be True and the name of your GPU.
Installing TensorFlow with CUDA Support
For TensorFlow, version 2.15 and later supports CUDA 12 by default:
pip install tensorflow
Verify the installation:
python -c "import tensorflow as tf; print(tf.config.list_physical_devices('GPU'))"
Running Your First Workload
To ensure everything is working correctly, run a simple workload. Create a file named test_gpu.py with the following content:
import torch
import time
# Check GPU availability
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
print(f"Using device: {device}")
# Create a large tensor and perform matrix operations
a = torch.randn(10000, 10000, device=device)
b = torch.randn(10000, 10000, device=device)
start = time.time()
c = torch.matmul(a, b)
torch.cuda.synchronize()
end = time.time()
print(f"Matrix multiplication time: {end - start:.4f} seconds")
print(f"Result shape: {c.shape}")
Run it:
python test_gpu.py
If everything is correct, the output will include the computation time and the matrix shape. This time should be significantly less than a similar run on the CPU.
Troubleshooting Common Issues
CUDA Out of Memory Error
This error occurs when the GPU memory is full. Solutions:
- Reduce the batch size during model training
- Use
torch.cuda.empty_cache()to free up cached memory - Use mixed precision training with
torch.cuda.amp
GPU Not Detected by PyTorch
If torch.cuda.is_available() returns False, the PyTorch version is likely incompatible with the installed CUDA. Check the PyTorch version:
python -c "import torch; print(torch.__version__)"
If the version ends with +cpu, it means the CPU version is installed. You need to install the CUDA version.
Reduced GPU Performance
If the GPU is working but performance is low, check the following:
- GPU temperature: Check the temperature with
nvidia-smi. If it's above 85°C, improve cooling. - Power consumption: Make sure the GPU is in maximum performance mode:
nvidia-smi -pm 1 - PCIe Gen4 usage: If your motherboard and GPU support PCIe Gen4, make sure it's enabled in the BIOS.
Conclusion
Setting up a GPU server for AI models is a multi-step process that requires precision in installing the driver, CUDA, and cuDNN. By following the steps in this article, you can create a stable and optimized environment for training and running your models. Always remember to keep the CUDA and driver versions compatible with each other, and before running heavy projects, verify GPU performance with a test workload. If you're looking for cloud infrastructure with GPU, ServerNet cloud services can be a suitable option to get started, but the final choice depends on your computational needs and budget.