SDS Computing Server: User Guide and Best Practices

Department of Statistics and Data Sciences, The University of Texas at Austin

This guide is for students, postdocs, and faculty using the SDS high-performance computing server. It covers setup, daily usage, and the practices that keep the server running smoothly for everyone.

If you remember nothing else: 1. Connect to the UT VPN first, then ssh in. 2. Run every job with nice, inside a tmux or screen session. 3. Never use all the cores — cap each job at 8–16 and set the thread-limit variables. 4. Check htop and nvidia-smi before starting anything heavy. 5. Clean up after yourself: kill finished processes, delete old files, environments, and installed software.


Quick Reference Cheat Sheet

The one-page version. Full explanations are in the sections below.

Connection

# VPN on first (on AND off campus), then:
ssh UTEID@CNS-SRV-SDS01.austin.utexas.edu
ssh sds            # if alias configured

Running jobs

# Quick job
nice python script.py

# Job that survives logout
tmux new -s myjob
nice python script.py > out.log 2>&1
# Ctrl+b, then d  to detach
# (avoid nohup — unreliable here)

Resource checks (before you run)

htop               # CPU + memory (Shift+H)
nvidia-smi         # GPU status
free -h            # RAM usage
df -h              # disk space
du -sh ~/          # your home size
ps aux --sort=-%cpu | head   # top CPU
ps aux --sort=-%mem | head   # top RAM
# per-user totals: see Monitoring section

Limit your usage

# put in ~/.bashrc
export OMP_NUM_THREADS=8
export MKL_NUM_THREADS=8
export OPENBLAS_NUM_THREADS=8

GPU selection

CUDA_VISIBLE_DEVICES=0 python train.py
CUDA_VISIBLE_DEVICES=2,3 python train.py

File transfer

scp file.py sds:~/projects/    # upload
scp sds:~/out.csv ./           # download
rsync -avz ./proj/ sds:~/proj/ # sync

Cleanup

conda env remove -n old_env
rm -r ~/old_project/
tmux kill-session -t old_sess

1. What Is the SDS Server?

The SDS server is a shared machine available to everyone in the department. It is designed for:

  • High-capacity statistical simulations (e.g., parallel MCMC, bootstrap)
  • Training deep learning models
  • Any computationally intensive work that exceeds what your laptop can handle

Hardware

Component Specification
CPU 2x AMD EPYC 7763 — 128 physical cores, 256 threads (255 online)
GPU 4x NVIDIA RTX A4000 (16 GB memory each)
RAM ~881 GiB usable, shared across all users
Storage Modest — treat it as a workspace, not an archive
OS Ubuntu 20.04 LTS

CPU and memory figures verified on the server August 13, 2026 via lscpu and free -h. GPU specification is from the department’s server page and has not been re-verified.

Key thing to understand

This is a shared machine with no job scheduler. Unlike large university clusters (e.g., TACC) that use schedulers like SLURM to allocate resources, the SDS server gives you direct access. This makes it easy to use but also easy to accidentally monopolize. When you take all 128 cores, everyone else’s work grinds to a halt. The best practices in this guide exist to prevent that.

Note that the machine reports 255 online logical CPUs, but that is 128 physical cores with two hardware threads each. Tools like os.cpu_count() in Python and detectCores() in R return 255. Treat 128 as the real number of cores, and never pass 255 to anything.


2. Connecting to the Server

Step 1: Connect to the UT VPN

You must be connected to the UT VPN to reach the server: on campus and off. Plain campus Wi-Fi is not enough and the server is not reachable without the VPN. Always connect it first, using the Cisco AnyConnect client.

  1. Download and install Cisco AnyConnect if you haven’t already.
  2. Open Cisco AnyConnect and connect to vpn.utexas.edu.
  3. Log in with your UT EID and password.
  4. For the Duo prompt, type push to receive a push notification on your phone, then approve it.
  5. You are connected when you see the locked lock symbol.

Stuck at a “Potential CSRF attack detected” page during login? That comes from stale cookies in the browser-based VPN/Duo login, not from the server. Clear your browser’s cache and cookies (or open the login in a private window) and connect again.

Step 2: SSH into the server

Open your terminal (Terminal on Mac, or Windows Terminal / PowerShell on Windows) and run:

ssh UTEID@CNS-SRV-SDS01.austin.utexas.edu

Replace UTEID with your actual UT EID (e.g., abc1234). Enter your password when prompted. You’ll see a welcome message if it worked.

Setting up SSH key authentication (optional)

To avoid typing your password every time:

  1. On your local machine, generate a key pair (if you don’t already have one):

    ssh-keygen -t ed25519

    Press Enter to accept defaults. You can set a passphrase or leave it empty.

  2. Copy the public key to the server:

    ssh-copy-id UTEID@CNS-SRV-SDS01.austin.utexas.edu
  3. Now ssh sds will log you in without a password (VPN still required).


4. Transferring Files

You have two main options for moving files between your computer and the server.

Option A: Cyberduck (graphical, beginner-friendly)

Cyberduck is a free file transfer app with a drag-and-drop interface.

  1. Open Cyberduck and go to File > Open Connection.
  2. Select SFTP (SSH File Transfer Protocol).
  3. Enter:
    • Server: CNS-SRV-SDS01.austin.utexas.edu
    • Port: 22
    • Username: Your UT EID
    • Password: Your UT password
  4. Click Connect.
  5. You can now drag and drop files between your computer and the server.

Option B: scp from the terminal

Copy a file to the server:

scp myfile.py sds:~/projects/

Copy a file from the server:

scp sds:~/results/output.csv ./

Copy an entire folder (use -r for recursive):

scp -r sds:~/results/ ./local_results/

5. Setting Up Your Environment

Python (Conda)

Do not use the system Python. Always use a Conda environment. This keeps your packages isolated and prevents conflicts with other users.

Installing Conda (first time only)

If Conda is not already installed in your home directory:

# Download Miniforge (lightweight Conda installer)
curl -L -O "https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-$(uname)-$(uname -m).sh"

# Run the installer
bash Miniforge3-$(uname)-$(uname -m).sh

Follow the prompts, accept the defaults, and say yes when asked to initialize Conda. Then close and reopen your terminal (or log out and back in).

You should now see (base) at the start of your prompt.

Creating an environment

Create a separate environment for each project:

# Create an environment with a specific Python version
conda create -n myproject python=3.11

# Activate it
conda activate myproject

# Install packages
pip install numpy pandas scikit-learn
# or
conda install numpy pandas scikit-learn

For deep learning (PyTorch)

conda create -n torch_env python=3.11
conda activate torch_env
pip install torch torchvision torchaudio

Verify PyTorch sees the GPUs:

import torch
print(torch.__version__)
print(f"CUDA available: {torch.cuda.is_available()}")
print(f"GPUs: {torch.cuda.device_count()}")

Listing and managing environments

conda env list              # See all your environments
conda activate myproject    # Switch to an environment
conda deactivate            # Return to base
conda env remove -n old_env # Delete an environment you no longer need

R

R is available on the server. To start an interactive R session:

R

Installing R packages

Inside the R console:

install.packages("tidyverse")
install.packages("data.table")

R will install packages to a user-local library in your home directory, so this doesn’t require admin permissions.

When you’re done, quit R:

q()

It will ask whether to save the workspace. Type y or n depending on your preference.

6. Running Jobs

Interactive vs. background jobs

  • Interactive: You run a command and wait for it to finish. Fine for quick tests (under a few minutes). If you close your terminal, the job dies.
  • Background: The job runs independently of your terminal session. Use this for anything that takes more than a few minutes.

Running a quick interactive job

Python:

conda activate myproject
python my_script.py

R:

R CMD BATCH my_script.R

This creates my_script.Rout with the output.

Running background jobs with nice

Always use nice when running background jobs. The nice command lowers the priority of your job so that it doesn’t starve other users’ work. This is a basic courtesy on a shared machine.

Python:

nice python my_script.py > output.log 2>&1 &

R:

nice R CMD BATCH my_script.R &

The & at the end sends the job to the background. You can check on it with:

jobs

Important: a job started this way is tied to your terminal session — it may die when you close the terminal or lose your VPN connection. For anything that runs longer than a few minutes, start it inside tmux or screen (next section).

Keeping jobs alive after logout: tmux

When you close your terminal or lose your VPN connection, background jobs started with & alone may die. Use tmux (or screen) to keep them running.

tmux basics

# Start a new named session
tmux new -s myjob

# You're now inside a tmux session. Run your job:
nice python train.py > train.log 2>&1

# Detach from the session (job keeps running):
# Press Ctrl+b, then d

# Log out, close terminal, disconnect VPN — the job survives.

# Later, reconnect and reattach:
tmux attach -t myjob

# List all tmux sessions:
tmux ls

# Kill a session when you're done:
tmux kill-session -t myjob

screen basics (alternative)

# Start a screen session
screen

# Run your job
nice R CMD BATCH my_sim.R &

# Detach: press Ctrl+a, then d

# Reattach later:
screen -r

# If multiple screens exist, list them:
screen -ls

# Reattach to a specific one:
screen -r SESSION_ID

A warning about nohup

On most Linux servers, nohup command & is the standard way to keep a job alive after logout. On this server, nohup has proven unreliable — jobs started with nohup have been observed to die anyway when the user logs out. Use tmux or screen instead; they are the dependable options here.

If you do try nohup, verify that your job actually survived: log out, log back in, and check that it’s still running:

ps -u $USER -f | grep python

7. Using GPUs

The server has 4 NVIDIA RTX A4000 GPUs, each with 16 GB of memory.

Check GPU availability before using one

Always run this before starting a GPU job:

nvidia-smi

This shows you which GPUs are in use, by whom, and how much memory is taken. If a GPU is already heavily loaded, pick a different one.

Selecting a specific GPU

By default, frameworks like PyTorch see all 4 GPUs. To restrict your job to specific GPUs, set the CUDA_VISIBLE_DEVICES environment variable.

Use GPU 0 only:

CUDA_VISIBLE_DEVICES=0 python train.py

Use GPUs 2 and 3:

CUDA_VISIBLE_DEVICES=2,3 python train.py

In Python, you can also set this programmatically (must be done before importing PyTorch/TensorFlow):

import os
os.environ["CUDA_VISIBLE_DEVICES"] = "1"

import torch  # Now only sees GPU 1

GPU usage in R

For R packages that support GPU (e.g., torch for R, keras/tensorflow):

Sys.setenv(CUDA_VISIBLE_DEVICES = "0")
library(torch)

GPU etiquette

  • Check nvidia-smi first. Don’t blindly grab all 4 GPUs.
  • Use only what you need. Most training jobs need 1 GPU. Only use multiple if your code explicitly supports multi-GPU training.
  • Release GPUs when done. Don’t leave idle processes sitting on GPUs. Kill finished or crashed processes.
  • Monitor GPU memory. If your job runs out of GPU memory, reduce batch size rather than grabbing another GPU.

8. Using Jupyter Notebook Remotely

You can run Jupyter Notebook on the server and access it from your local browser.

Step 1: Start Jupyter on the server

SSH into the server and run:

conda activate myproject
jupyter notebook --no-browser --port=9999

Jupyter will print a URL with a token. Keep this terminal open (or use tmux).

Step 2: Create an SSH tunnel from your local machine

Open a new terminal window on your local machine and run:

ssh -NfL 9999:localhost:9999 sds

(If you didn’t set up the SSH alias, use UTEID@CNS-SRV-SDS01.austin.utexas.edu instead of sds.)

Step 3: Open in your browser

Go to http://localhost:9999 in your browser. Paste the token from Step 1 when prompted.

Picking a port

If port 9999 is already taken by another user, pick a different number (e.g., 8888, 8877, or any number between 1024-65535). Use the same port in both the Jupyter command and the SSH tunnel.

Cleaning up

When you’re done, stop the Jupyter server with Ctrl+C in the server terminal. To kill the SSH tunnel on your local machine:

# Find the tunnel process
ps aux | grep "ssh -NfL"

# Kill it by PID
kill PID

9. Monitoring Your Jobs and System Resources

Check who and what is running

# See your own running processes
ps -u $USER -f

Figuring out who is using resources

On a shared machine with no scheduler, sometimes the answer to “why is my job slow?” is that someone else is using the resources. Here is how to find out who — not to police anyone, but so you can coordinate. A quick email or message usually solves it.

Who is logged in right now:

w

Top CPU and memory consumers, with their owners:

ps -eo user,pid,%cpu,%mem,etime,args --sort=-%cpu | grep -E '[p]ython|[R] ' | head -15

The [p]ython bracket trick keeps grep from matching its own command line. Typical output (user IDs are made up):

USER         PID %CPU %MEM   ELAPSED COMMAND
ab12345  1080772  863  0.1  18:29:00 /usr/lib/R/bin/exec/R -f simulations_scenario1.R
ab12345  1184147  407  1.1  18:00:29 /usr/lib/R/bin/exec/R -f simulations_scenario2.R
cd67890  3216392  103  0.1     00:13 /usr/lib/R/bin/exec/R -f run_setting6.R
cd67890  3216351  102  0.1     00:14 /usr/lib/R/bin/exec/R -f run_setting6.R

Reading the %CPU column. %CPU is a process’s CPU time divided by its wall-clock lifetime, summed across cores — so 100% = one core kept fully busy. In the example, 863 means the job is using about 8.6 cores; 102 means about one. Two things to keep in mind: it is a lifetime average, not a live reading, so a job that just started (small ELAPSED) has a less settled number; and on this 128-core machine a single multi-threaded job can legitimately show several thousand percent. For an instantaneous, live view use htop or top.

Total usage per user:

ps -eo user,%cpu,%mem --no-headers | awk '{cpu[$1]+=$2; mem[$1]+=$3} END {for (u in cpu) printf "%-12s %7.1f%% CPU %6.1f%% MEM\n", u, cpu[u], mem[u]}' | sort -k2 -rn

This sums CPU and memory percentages across each user’s processes. (CPU percentages are per-core, so with 255 online CPUs the totals across all users can approach 25500%.)

Memory per user, in gigabytes:

Percentages are hard to reason about. This reports actual memory per user:

ps -eo user,rss --no-headers | awk '{m[$1]+=$2} END {for (u in m) printf "%-14s %8.1f GiB\n", u, m[u]/1048576}' | sort -k2 -rn

Caveat: this sums RSS, which double-counts memory shared between a parent process and its forked workers. It is reliable for identifying who is using a lot, not for exact accounting. For true per-user figures without double-counting, use systemd-cgtop -m --depth=2 and read the user.slice/user-NNNN.slice rows (getent passwd NNNN maps the UID to a name).

Biggest individual processes, with elapsed time and full command:

ps -eo pid,user,rss,%mem,etime,args --sort=-rss | head -20

rss is in kilobytes. etime tells you whether something is a long-running job or just started, which is useful context before you contact anyone.

Who is using the GPUs:

nvidia-smi lists the PIDs of GPU processes but not who owns them. Feed those PIDs to ps to get usernames and full commands:

ps -up $(nvidia-smi --query-compute-apps=pid --format=csv,noheader)

If no processes are running on any GPU, the PID list is empty and ps will complain — that just means the GPUs are free.

# For Full GPU status
nvidia-smi

# Compact view
nvidia-smi --query-gpu=index,name,memory.used,memory.total,utilization.gpu --format=csv,noheader

# Auto-refreshing GPU monitor (updates every 2 seconds)
watch -n 2 nvidia-smi

What to do with this information: if someone is using most of the machine, don’t kill anything, and don’t just wait silently either. Email them, ask in the department channel, or contact the admin. Most of the time the person doesn’t realize their job grabbed everything, and a quick message fixes it.

Check disk space

# Overall disk usage
df -h

# Your home directory size
du -sh ~/

# Find what's taking the most space in your home directory
du -sh ~/* | sort -rh | head -20

Check your running jobs

# If you used &
jobs

# Your Python processes specifically
ps -u $USER -f | grep python

# Your tmux sessions
tmux ls

# Your screen sessions
screen -ls

10. Best Practices: How Not to Crash the Server

This is the most important section of this guide. The SDS server is a shared resource with no job scheduler, which means a single user’s runaway job can bring the entire server down for everyone.

Rule 1: Always nice your jobs

The nice command tells the operating system to give your job lower scheduling priority. This means if someone is running an interactive session, your batch job won’t freeze them out.

# Good
nice python my_simulation.py

# Bad — runs at normal priority, hogs resources
python my_simulation.py

For even lower priority:

nice -n 10 python my_simulation.py

Rule 2: Limit your CPU cores

This is the single most common way the server gets crashed. Many libraries in Python and R will, by default, use every available CPU core. On this machine os.cpu_count() and detectCores() both return 255, so a single job left at its defaults can spawn 255 workers and consume the entire server.

Python: Control parallel workers

NumPy, SciPy, and other libraries using BLAS/OpenMP:

These libraries use multi-threaded linear algebra under the hood. Limit them with environment variables:

export OMP_NUM_THREADS=8
export MKL_NUM_THREADS=8
export OPENBLAS_NUM_THREADS=8

Put these in your ~/.bashrc so they apply every time you log in:

echo 'export OMP_NUM_THREADS=8' >> ~/.bashrc
echo 'export MKL_NUM_THREADS=8' >> ~/.bashrc
echo 'export OPENBLAS_NUM_THREADS=8' >> ~/.bashrc
source ~/.bashrc

PyTorch DataLoader:

# Bad — os.cpu_count() returns 255 on this machine, so this spawns 255 workers
DataLoader(dataset, num_workers=os.cpu_count())

# Good — use a reasonable number
DataLoader(dataset, num_workers=4)

multiprocessing and joblib:

# Bad
from multiprocessing import Pool
pool = Pool()  # Defaults to ALL cores

# Good
pool = Pool(8)  # Use 8 cores

# With joblib
from joblib import parallel_backend
with parallel_backend('loky', n_jobs=8):
    results = Parallel()(delayed(func)(x) for x in data)

Pandas:

Some pandas operations are parallelized internally. Control this with the same OMP_NUM_THREADS variable above.

R: Control parallel cores

parallel package:

# Bad — detectCores() returns 255 on this machine
library(parallel)
mclapply(data, func, mc.cores = detectCores())

# Good — leave cores for others
mclapply(data, func, mc.cores = 8)

foreach and doParallel:

# Bad
library(doParallel)
registerDoParallel(cores = detectCores())

# Good
registerDoParallel(cores = 8)

BLAS threads (R uses these for matrix operations):

Add to your ~/.Renviron file:

OMP_NUM_THREADS=8
MKL_NUM_THREADS=8
OPENBLAS_NUM_THREADS=8

General rule of thumb: Use no more than 8-16 physical cores per job unless you have confirmed that no one else is running anything. Check with htop before deciding.

Rule 3: Watch your memory usage

The server has roughly 881 GiB of RAM, shared across all users. Loading very large datasets entirely into memory, or running many parallel workers each with their own copy of data, can exhaust memory and take the machine down for everyone.

  • Load only the data you need (use chunking, lazy loading, or memory-mapped files)
  • Monitor with free -h and htop
  • If your job needs more than ~150 GB of RAM, check that enough is available first

Memory is the one you can actually break. Oversubscribing the CPUs makes everyone slow, which is annoying but recoverable. Exhausting memory is worse: when the machine runs out, the Linux “OOM killer” starts terminating processes to recover — and it does not necessarily pick yours. Someone else’s three-day simulation can die because of your data frame. Before a large job, check free -h and stay well within what is free.

Python example — memory-mapped files for large data:

import numpy as np
# Instead of loading everything into RAM:
data = np.load("huge_array.npy", mmap_mode='r')

R example — chunked reading:

library(data.table)
# Instead of read.csv (loads everything at once):
dt <- fread("large_file.csv", nrows = 100000)  # Read first 100k rows

Rule 4: Keep your home directory clean

Storage on the server is limited and shared. Large files from one user can fill up the disk and cause problems for everyone, including the system itself.

  • Delete data you no longer need — old datasets, intermediate results, log files
  • Delete Conda environments you’re done with: conda env remove -n old_env
  • Delete installed software you no longer use — this is explicitly called out in the server usage policy
  • Don’t store large datasets permanently — transfer results to your local machine or UT’s research storage

Monitor your disk space

Check the whole disk and your own footprint:

df -h        # usage of every filesystem — look at the one holding /home
du -sh ~/    # total size of your home directory

Find what is eating the space

List each folder in your home directory, biggest first. Include hidden folders like .cache and .conda — on this server they are usually the real culprits:

# Top 10 folders, visible and hidden, largest first
du -sh ~/.[!.]* ~/* 2>/dev/null | sort -rh | head -10

The 2>/dev/null silences “permission denied” noise. The biggest item is almost always ~/miniforge3 (your Conda install), followed by dataset and checkpoint folders.

Shrink your Conda / miniforge footprint

A miniforge install easily grows to tens of GB. First see where the space goes:

du -sh ~/miniforge3/envs/* | sort -rh   # size of each environment
conda env list                          # environments you still have

Then reclaim space in two safe steps:

# 1. Clear Conda's download + package cache
conda clean --all

# 2. Delete whole environments you no longer use (the biggest win)
conda env remove -n old_env

Is conda clean safe? Yes. conda clean --all asks separately about tarballs, the index cache, and packages. All three act only on the package cache in ~/miniforge3/pkgs — downloaded archives and unpacked entries that no environment is currently using. It does not remove packages from your active environments, so answering y to all three is safe. The only cost is that reinstalling a cleaned package later re-downloads it. Deleting an entire unused environment with conda env remove frees far more than cache cleaning ever will.

Rule 5: Don’t hog the GPUs

  • Run nvidia-smi before starting a GPU job
  • Set CUDA_VISIBLE_DEVICES to use only the GPUs you need (usually 1)
  • Kill finished or crashed GPU processes promptly
  • Don’t leave Jupyter notebooks with loaded GPU models sitting idle

Rule 6: Clean up after yourself

  • Kill background processes when they’re done or if they’ve failed
  • Close tmux/screen sessions you’re no longer using
  • Remove large temporary files, checkpoints, and log files you don’t need
  • Deactivate and remove unused Conda environments

Summary of resource limits to follow

These numbers are community etiquette, not enforced limits — the server has no scheduler to stop you, and they are not official policy. When in doubt (e.g., you genuinely need more for a deadline), check with the server administrator or coordinate with other users first.

Resource Recommended Limit
CPU cores per job 8-16 physical cores (check htop first)
Total CPU cores per user No more than ~32 unless the server is idle
RAM per job Check free -h; stay under ~150 GB
GPUs per user 1-2 (check nvidia-smi first)
PyTorch num_workers 4-8 — never os.cpu_count(), which returns 255 here
Home directory size Clean up regularly; no permanent bulk storage

11. Troubleshooting

“Connection refused” or “Connection timed out”

  • Make sure your VPN is connected
  • Verify the hostname: CNS-SRV-SDS01.austin.utexas.edu
  • Try pinging the server: ping CNS-SRV-SDS01.austin.utexas.edu

“Permission denied” on SSH

  • Double-check your EID and password
  • If using SSH keys, verify the key is loaded: ssh-add -l

My job died when I closed the terminal

  • You need to use tmux or screen (see Running Jobs)
  • tmux new -s myjob before running your job, then Ctrl+b, d to detach
  • Note: nohup is not reliable on this server — jobs started with it may still die at logout

My Python/R script can’t find a package

  • Make sure you activated the right Conda environment: conda activate myproject
  • For R, check if the package is installed: installed.packages()
  • Install missing packages within your environment, not globally

nvidia-smi shows my GPU process but the job crashed

The process might be a zombie. Kill it:

# Find your GPU processes
nvidia-smi

# Kill by PID
kill PID

# If it won't die
kill -9 PID

The server feels extremely slow

Someone (maybe you) might be using too many cores. Check:

htop

If the CPU meters are all pegged and the load average is well above 128, find the top consumers and who owns them:

ps aux --sort=-%cpu | head -15

See Figuring out who is using resources for more ways to break down usage by user. Then talk to the person (or email the admin) — don’t kill anyone else’s processes.

The GPUs are all busy — who is using them?

nvidia-smi shows the PIDs of GPU processes but not their owners. Map the PIDs to usernames and commands:

ps -up $(nvidia-smi --query-compute-apps=pid --format=csv,noheader)

Then coordinate with that person directly, or wait for their job to finish — never kill another user’s process.

“No space left on device”

The disk is full. Check your usage:

du -sh ~/

Clean up what you can. If you’re not the cause, contact the admin.

CUDA out of memory

Reduce your batch size, or use a different (less occupied) GPU:

CUDA_VISIBLE_DEVICES=2 python train.py --batch_size 16

For questions or issues, contact the SDS server administrator. For suggested edits to this guide, reach out to Ritwik Vashistha.