Setting up a GPU server for AI models

Step-by-step guide to setting up a GPU server for AI models: installing drivers, CUDA, cuDNN, and running your first workload. Includes practical examples and troubleshooting common errors.

6 min Updated 30 Sep 2026

Why Is Setting Up a GPU Server for AI Models Critical?

Running deep learning models on a CPU is not only slow, but it becomes practically impossible for modern models with billions of parameters. A well-equipped GPU server can reduce the training time of a model from several weeks to a few hours. However, purchasing or renting hardware is only part of the story; correctly installing the driver, CUDA Toolkit, and related libraries determines your success. In this article, we will follow the complete path of setting up a GPU server from the moment you have access to the system to running your first workload, with technical details and real-world examples.

Prerequisites and Hardware Check

Before taking any action, you must ensure that your GPU is recognized by the operating system and has the appropriate driver. If you are using a cloud server (for example, ServerNet cloud services that offer GPU server rentals), you will typically receive a pre-configured image. But if you have a dedicated server, follow the steps below.

Identifying the GPU in Linux

Most GPU servers run Ubuntu 20.04 or 22.04 LTS. First, check whether the GPU is detected by the system using the following command:

lspci | grep -i nvidia

The output should look something like this:

01:00.0 3D controller: NVIDIA Corporation GA102 [GeForce RTX 3080] (rev a1)

If the output is empty, the GPU is likely not properly seated in the PCIe slot or is disabled in the BIOS. On cloud servers, this step is usually already done.

Checking the Current Driver

To see the installed driver version (if any), use the following command:

nvidia-smi

If this command does not work, the driver is not installed or not properly configured. A successful output should include the GPU name, driver version, and CUDA version (which usually comes with the driver).

Common Mistake: Many users think installing the CUDA Toolkit alone is sufficient. But the CUDA Toolkit does not work without the appropriate driver. The driver and CUDA are two separate layers that must be compatible with each other.

Installing the NVIDIA Driver

There are two main methods for installing the driver: using the official NVIDIA repository or the Ubuntu repository. The first method is recommended because it provides a newer version.

Method 1: Installing from the Official NVIDIA Repository

First, add the official repository:

sudo add-apt-repository ppa:graphics-drivers/ppa
sudo apt update

Then install the appropriate driver version. To find the latest stable version, you can search:

apt search nvidia-driver | grep -E '^nvidia-driver-[0-9]+'

The output usually includes versions like nvidia-driver-535 or nvidia-driver-545. Install version 535 or higher:

sudo apt install nvidia-driver-535

After installation, reboot the system:

sudo reboot

Method 2: Installing from the Ubuntu Repository

If the above method fails for any reason, you can use the default repository:

sudo apt install nvidia-driver-535-server

This version is optimized for servers and has greater stability.

Verifying the Driver Installation

After rebooting, run the nvidia-smi command again. The output should include a table with GPU specifications and the driver version. If you receive the error NVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver, the driver has not been loaded correctly. In this case, check the logs:

dmesg | grep -i nvidia

Troubleshooting Tip: If Secure Boot is enabled, the driver may be unsigned and fail to load. Disable Secure Boot in the BIOS or sign the driver with the MOK key.

Installing the CUDA Toolkit

The CUDA Toolkit includes the nvcc compiler, runtime libraries, and development tools. The CUDA version must be compatible with the driver version. For example, driver 535 supports CUDA 12.2.

Downloading and Installing CUDA

Visit the official CUDA Toolkit page and select the appropriate version for your operating system. For Ubuntu 22.04, run the following commands:

wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt update
sudo apt install cuda-toolkit-12-2

After installation, add the CUDA paths to the environment variables. Do this by editing the ~/.bashrc file:

echo 'export PATH=/usr/local/cuda-12.2/bin:$PATH' >> ~/.bashrc
echo 'export LD_LIBRARY_PATH=/usr/local/cuda-12.2/lib64:$LD_LIBRARY_PATH' >> ~/.bashrc
source ~/.bashrc

Verifying the CUDA Installation

Check the CUDA version:

nvcc --version

The output should include a version like Cuda compilation tools, release 12.2, V12.2.140.

Installing cuDNN for Optimization

cuDNN (CUDA Deep Neural Network library) is an accelerator for neural network operations. Without it, frameworks like TensorFlow and PyTorch use slower implementations.

Installing cuDNN

To install cuDNN, you need to register on the NVIDIA website and download the appropriate files. After downloading, follow these steps:

tar -xzvf cudnn-linux-x86_64-8.9.7.29_cuda12-archive.tar.xz
sudo cp cudnn-linux-x86_64-8.9.7.29_cuda12-archive/include/cudnn*.h /usr/local/cuda-12.2/include/
sudo cp -P cudnn-linux-x86_64-8.9.7.29_cuda12-archive/lib/libcudnn* /usr/local/cuda-12.2/lib64/
sudo chmod a+r /usr/local/cuda-12.2/include/cudnn*.h /usr/local/cuda-12.2/lib64/libcudnn*

To verify the installation, check the cuDNN version:

cat /usr/local/cuda-12.2/include/cudnn_version.h | grep CUDNN_MAJOR -A 2

Installing Deep Learning Frameworks

Now that the driver, CUDA, and cuDNN are ready, it's time to install the framework. PyTorch and TensorFlow are the two main options.

Installing PyTorch with CUDA Support

To install PyTorch with CUDA 12.1 (which is compatible with CUDA 12.2), use the following command:

pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121

Verify that PyTorch can see the GPU:

python -c "import torch; print(torch.cuda.is_available()); print(torch.cuda.get_device_name(0))"

The output should be True and the name of your GPU.

Installing TensorFlow with CUDA Support

For TensorFlow, version 2.15 and later supports CUDA 12 by default:

pip install tensorflow

Verify the installation:

python -c "import tensorflow as tf; print(tf.config.list_physical_devices('GPU'))"

Running Your First Workload

To ensure everything is working correctly, run a simple workload. Create a file named test_gpu.py with the following content:

import torch
import time

# Check GPU availability
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
print(f"Using device: {device}")

# Create a large tensor and perform matrix operations
a = torch.randn(10000, 10000, device=device)
b = torch.randn(10000, 10000, device=device)

start = time.time()
c = torch.matmul(a, b)
torch.cuda.synchronize()
end = time.time()

print(f"Matrix multiplication time: {end - start:.4f} seconds")
print(f"Result shape: {c.shape}")

Run it:

python test_gpu.py

If everything is correct, the output will include the computation time and the matrix shape. This time should be significantly less than a similar run on the CPU.

Troubleshooting Common Issues

CUDA Out of Memory Error

This error occurs when the GPU memory is full. Solutions:

  • Reduce the batch size during model training
  • Use torch.cuda.empty_cache() to free up cached memory
  • Use mixed precision training with torch.cuda.amp

GPU Not Detected by PyTorch

If torch.cuda.is_available() returns False, the PyTorch version is likely incompatible with the installed CUDA. Check the PyTorch version:

python -c "import torch; print(torch.__version__)"

If the version ends with +cpu, it means the CPU version is installed. You need to install the CUDA version.

Reduced GPU Performance

If the GPU is working but performance is low, check the following:

  • GPU temperature: Check the temperature with nvidia-smi. If it's above 85°C, improve cooling.
  • Power consumption: Make sure the GPU is in maximum performance mode: nvidia-smi -pm 1
  • PCIe Gen4 usage: If your motherboard and GPU support PCIe Gen4, make sure it's enabled in the BIOS.

Conclusion

Setting up a GPU server for AI models is a multi-step process that requires precision in installing the driver, CUDA, and cuDNN. By following the steps in this article, you can create a stable and optimized environment for training and running your models. Always remember to keep the CUDA and driver versions compatible with each other, and before running heavy projects, verify GPU performance with a test workload. If you're looking for cloud infrastructure with GPU, ServerNet cloud services can be a suitable option to get started, but the final choice depends on your computational needs and budget.

Was this page helpful?