Department of Statistics and Data Sciences, The University of Texas at Austin
This guide is for students, postdocs, and faculty using the SDS high-performance computing server. It covers setup, daily usage, and the practices that keep the server running smoothly for everyone.
If you remember nothing else: 1. Connect to the UT VPN first, then
sshin. 2. Run every job withnice, inside atmuxorscreensession. 3. Never use all the cores — cap each job at 8–16 and set the thread-limit variables. 4. Checkhtopandnvidia-smibefore starting anything heavy. 5. Clean up after yourself: kill finished processes, delete old files, environments, and installed software.
The one-page version. Full explanations are in the sections below.
Connection
# VPN on first (on AND off campus), then:
ssh UTEID@CNS-SRV-SDS01.austin.utexas.edu
ssh sds # if alias configured
Running jobs
# Quick job
nice python script.py
# Job that survives logout
tmux new -s myjob
nice python script.py > out.log 2>&1
# Ctrl+b, then d to detach
# (avoid nohup — unreliable here)
Resource checks (before you run)
htop # CPU + memory (Shift+H)
nvidia-smi # GPU status
free -h # RAM usage
df -h # disk space
du -sh ~/ # your home size
ps aux --sort=-%cpu | head # top CPU
ps aux --sort=-%mem | head # top RAM
# per-user totals: see Monitoring section
Limit your usage
# put in ~/.bashrc
export OMP_NUM_THREADS=8
export MKL_NUM_THREADS=8
export OPENBLAS_NUM_THREADS=8
GPU selection
CUDA_VISIBLE_DEVICES=0 python train.py
CUDA_VISIBLE_DEVICES=2,3 python train.py
File transfer
scp file.py sds:~/projects/ # upload
scp sds:~/out.csv ./ # download
rsync -avz ./proj/ sds:~/proj/ # sync
Cleanup
conda env remove -n old_env
rm -r ~/old_project/
tmux kill-session -t old_sess
The SDS server is a shared machine available to everyone in the department. It is designed for:
| Component | Specification |
|---|---|
| CPU | 2x AMD EPYC 7763 — 128 physical cores, 256 threads (255 online) |
| GPU | 4x NVIDIA RTX A4000 (16 GB memory each) |
| RAM | ~881 GiB usable, shared across all users |
| Storage | Modest — treat it as a workspace, not an archive |
| OS | Ubuntu 20.04 LTS |
CPU and memory figures verified on the server August 13, 2026 via
lscpu and free -h. GPU specification is from
the department’s server page and has not been re-verified.
This is a shared machine with no job scheduler. Unlike large university clusters (e.g., TACC) that use schedulers like SLURM to allocate resources, the SDS server gives you direct access. This makes it easy to use but also easy to accidentally monopolize. When you take all 128 cores, everyone else’s work grinds to a halt. The best practices in this guide exist to prevent that.
Note that the machine reports 255 online logical
CPUs, but that is 128 physical cores with two hardware threads
each. Tools like os.cpu_count() in Python and
detectCores() in R return 255. Treat 128 as the real number
of cores, and never pass 255 to anything.
You must be connected to the UT VPN to reach the server: on campus and off. Plain campus Wi-Fi is not enough and the server is not reachable without the VPN. Always connect it first, using the Cisco AnyConnect client.
vpn.utexas.edu.push to receive a push
notification on your phone, then approve it.Stuck at a “Potential CSRF attack detected” page during login? That comes from stale cookies in the browser-based VPN/Duo login, not from the server. Clear your browser’s cache and cookies (or open the login in a private window) and connect again.
Open your terminal (Terminal on Mac, or Windows Terminal / PowerShell on Windows) and run:
ssh UTEID@CNS-SRV-SDS01.austin.utexas.edu
Replace UTEID with your actual UT EID (e.g.,
abc1234). Enter your password when prompted. You’ll see a
welcome message if it worked.
Typing the full hostname every time is time consuming. You can create a shortcut by adding an entry to your SSH config file.
Open (or create) ~/.ssh/config on your local
machine and add:
Host sds
HostName CNS-SRV-SDS01.austin.utexas.edu
User UTEID
Replace UTEID with your EID. Now you can connect with
just:
ssh sds
Never made this file before? On Mac or Linux, run these on your local machine:
mkdir -p ~/.ssh && chmod 700 ~/.ssh # create the folder if needed
nano ~/.ssh/config # opens the file (empty if new)
Paste the Host sds block above, save with
Ctrl+O then Enter, exit with Ctrl+X, and
tighten the permissions:
chmod 600 ~/.ssh/config
On Windows, the file lives at
C:\Users\<you>\.ssh\config (no file extension). Open
it with notepad %USERPROFILE%\.ssh\config; if Notepad saves
it as config.txt, rename it to remove the
.txt. One config file can hold as many Host
entries as you like — add a new block for each server.
To avoid typing your password every time:
On your local machine, generate a key pair (if you don’t already have one):
ssh-keygen -t ed25519
Press Enter to accept defaults. You can set a passphrase or leave it empty.
Copy the public key to the server:
ssh-copy-id UTEID@CNS-SRV-SDS01.austin.utexas.eduNow ssh sds will log you in without a password (VPN
still required).
You have two main options for moving files between your computer and the server.
Cyberduck is a free file transfer app with a drag-and-drop interface.
CNS-SRV-SDS01.austin.utexas.eduscp from the terminalCopy a file to the server:
scp myfile.py sds:~/projects/
Copy a file from the server:
scp sds:~/results/output.csv ./
Copy an entire folder (use -r for recursive):
scp -r sds:~/results/ ./local_results/
Do not use the system Python. Always use a Conda environment. This keeps your packages isolated and prevents conflicts with other users.
If Conda is not already installed in your home directory:
# Download Miniforge (lightweight Conda installer)
curl -L -O "https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-$(uname)-$(uname -m).sh"
# Run the installer
bash Miniforge3-$(uname)-$(uname -m).sh
Follow the prompts, accept the defaults, and say yes when asked to initialize Conda. Then close and reopen your terminal (or log out and back in).
You should now see (base) at the start of your
prompt.
Create a separate environment for each project:
# Create an environment with a specific Python version
conda create -n myproject python=3.11
# Activate it
conda activate myproject
# Install packages
pip install numpy pandas scikit-learn
# or
conda install numpy pandas scikit-learn
conda create -n torch_env python=3.11
conda activate torch_env
pip install torch torchvision torchaudio
Verify PyTorch sees the GPUs:
import torch
print(torch.__version__)
print(f"CUDA available: {torch.cuda.is_available()}")
print(f"GPUs: {torch.cuda.device_count()}")
conda env list # See all your environments
conda activate myproject # Switch to an environment
conda deactivate # Return to base
conda env remove -n old_env # Delete an environment you no longer need
R is available on the server. To start an interactive R session:
R
Inside the R console:
install.packages("tidyverse")
install.packages("data.table")
R will install packages to a user-local library in your home directory, so this doesn’t require admin permissions.
When you’re done, quit R:
q()
It will ask whether to save the workspace. Type y or
n depending on your preference.
renv for project-level package management
(recommended)Similar to Python environments, renv isolates R package
versions per project:
install.packages("renv")
renv::init() # Initialize in your project directory
renv::snapshot() # Save the current state of packages
renv::restore() # Restore packages from the lockfile
renv gives each project its own private package library
plus a renv.lock file that records the exact version of
every package you use. That makes your project reproducible — you (or a
collaborator) can rebuild the same package set later — and keeps it
isolated from other users’ R packages on the shared machine. In short:
init() sets up the project library, snapshot()
writes the current versions into the lockfile, and
restore() reinstalls exactly those versions from the
lockfile (on another machine, or after an upgrade breaks something).
Full guide: https://rstudio.github.io/renv/.
Python:
conda activate myproject
python my_script.py
R:
R CMD BATCH my_script.R
This creates my_script.Rout with the output.
niceAlways use nice when running background
jobs. The nice command lowers the priority of your
job so that it doesn’t starve other users’ work. This is a basic
courtesy on a shared machine.
Python:
nice python my_script.py > output.log 2>&1 &
R:
nice R CMD BATCH my_script.R &
The & at the end sends the job to the background.
You can check on it with:
jobs
Important: a job started this way is tied to your
terminal session — it may die when you close the terminal or lose your
VPN connection. For anything that runs longer than a few minutes, start
it inside tmux or screen (next section).
tmuxWhen you close your terminal or lose your VPN connection, background
jobs started with & alone may die. Use
tmux (or screen) to keep them running.
tmux basics# Start a new named session
tmux new -s myjob
# You're now inside a tmux session. Run your job:
nice python train.py > train.log 2>&1
# Detach from the session (job keeps running):
# Press Ctrl+b, then d
# Log out, close terminal, disconnect VPN — the job survives.
# Later, reconnect and reattach:
tmux attach -t myjob
# List all tmux sessions:
tmux ls
# Kill a session when you're done:
tmux kill-session -t myjob
screen basics (alternative)# Start a screen session
screen
# Run your job
nice R CMD BATCH my_sim.R &
# Detach: press Ctrl+a, then d
# Reattach later:
screen -r
# If multiple screens exist, list them:
screen -ls
# Reattach to a specific one:
screen -r SESSION_ID
nohupOn most Linux servers, nohup command & is the
standard way to keep a job alive after logout. On this server,
nohup has proven unreliable — jobs started with
nohup have been observed to die anyway when the user logs
out. Use tmux or screen instead; they are the
dependable options here.
If you do try nohup, verify that your job actually
survived: log out, log back in, and check that it’s still running:
ps -u $USER -f | grep python
The server has 4 NVIDIA RTX A4000 GPUs, each with 16 GB of memory.
Always run this before starting a GPU job:
nvidia-smi
This shows you which GPUs are in use, by whom, and how much memory is taken. If a GPU is already heavily loaded, pick a different one.
By default, frameworks like PyTorch see all 4 GPUs. To restrict your
job to specific GPUs, set the CUDA_VISIBLE_DEVICES
environment variable.
Use GPU 0 only:
CUDA_VISIBLE_DEVICES=0 python train.py
Use GPUs 2 and 3:
CUDA_VISIBLE_DEVICES=2,3 python train.py
In Python, you can also set this programmatically (must be done before importing PyTorch/TensorFlow):
import os
os.environ["CUDA_VISIBLE_DEVICES"] = "1"
import torch # Now only sees GPU 1
For R packages that support GPU (e.g., torch for R,
keras/tensorflow):
Sys.setenv(CUDA_VISIBLE_DEVICES = "0")
library(torch)
nvidia-smi first. Don’t blindly
grab all 4 GPUs.You can run Jupyter Notebook on the server and access it from your local browser.
SSH into the server and run:
conda activate myproject
jupyter notebook --no-browser --port=9999
Jupyter will print a URL with a token. Keep this terminal open (or use tmux).
Open a new terminal window on your local machine and run:
ssh -NfL 9999:localhost:9999 sds
(If you didn’t set up the SSH alias, use
UTEID@CNS-SRV-SDS01.austin.utexas.edu instead of
sds.)
Go to http://localhost:9999 in your browser. Paste the
token from Step 1 when prompted.
If port 9999 is already taken by another user, pick a different number (e.g., 8888, 8877, or any number between 1024-65535). Use the same port in both the Jupyter command and the SSH tunnel.
When you’re done, stop the Jupyter server with Ctrl+C in
the server terminal. To kill the SSH tunnel on your local machine:
# Find the tunnel process
ps aux | grep "ssh -NfL"
# Kill it by PID
kill PID
# See your own running processes
ps -u $USER -f
On a shared machine with no scheduler, sometimes the answer to “why is my job slow?” is that someone else is using the resources. Here is how to find out who — not to police anyone, but so you can coordinate. A quick email or message usually solves it.
Who is logged in right now:
w
Top CPU and memory consumers, with their owners:
ps -eo user,pid,%cpu,%mem,etime,args --sort=-%cpu | grep -E '[p]ython|[R] ' | head -15
The [p]ython bracket trick keeps grep from
matching its own command line. Typical output (user IDs are made
up):
USER PID %CPU %MEM ELAPSED COMMAND
ab12345 1080772 863 0.1 18:29:00 /usr/lib/R/bin/exec/R -f simulations_scenario1.R
ab12345 1184147 407 1.1 18:00:29 /usr/lib/R/bin/exec/R -f simulations_scenario2.R
cd67890 3216392 103 0.1 00:13 /usr/lib/R/bin/exec/R -f run_setting6.R
cd67890 3216351 102 0.1 00:14 /usr/lib/R/bin/exec/R -f run_setting6.R
Reading the
%CPUcolumn.%CPUis a process’s CPU time divided by its wall-clock lifetime, summed across cores — so 100% = one core kept fully busy. In the example,863means the job is using about 8.6 cores;102means about one. Two things to keep in mind: it is a lifetime average, not a live reading, so a job that just started (smallELAPSED) has a less settled number; and on this 128-core machine a single multi-threaded job can legitimately show several thousand percent. For an instantaneous, live view usehtoportop.
Total usage per user:
ps -eo user,%cpu,%mem --no-headers | awk '{cpu[$1]+=$2; mem[$1]+=$3} END {for (u in cpu) printf "%-12s %7.1f%% CPU %6.1f%% MEM\n", u, cpu[u], mem[u]}' | sort -k2 -rn
This sums CPU and memory percentages across each user’s processes. (CPU percentages are per-core, so with 255 online CPUs the totals across all users can approach 25500%.)
Memory per user, in gigabytes:
Percentages are hard to reason about. This reports actual memory per user:
ps -eo user,rss --no-headers | awk '{m[$1]+=$2} END {for (u in m) printf "%-14s %8.1f GiB\n", u, m[u]/1048576}' | sort -k2 -rn
Caveat: this sums RSS, which double-counts memory shared between a
parent process and its forked workers. It is reliable for identifying
who is using a lot, not for exact accounting. For true per-user figures
without double-counting, use systemd-cgtop -m --depth=2 and
read the user.slice/user-NNNN.slice rows
(getent passwd NNNN maps the UID to a name).
Biggest individual processes, with elapsed time and full command:
ps -eo pid,user,rss,%mem,etime,args --sort=-rss | head -20
rss is in kilobytes. etime tells you
whether something is a long-running job or just started, which is useful
context before you contact anyone.
Who is using the GPUs:
nvidia-smi lists the PIDs of GPU processes but not who
owns them. Feed those PIDs to ps to get usernames and full
commands:
ps -up $(nvidia-smi --query-compute-apps=pid --format=csv,noheader)
If no processes are running on any GPU, the PID list is empty and
ps will complain — that just means the GPUs are free.
# For Full GPU status
nvidia-smi
# Compact view
nvidia-smi --query-gpu=index,name,memory.used,memory.total,utilization.gpu --format=csv,noheader
# Auto-refreshing GPU monitor (updates every 2 seconds)
watch -n 2 nvidia-smi
What to do with this information: if someone is using most of the machine, don’t kill anything, and don’t just wait silently either. Email them, ask in the department channel, or contact the admin. Most of the time the person doesn’t realize their job grabbed everything, and a quick message fixes it.
# Overall disk usage
df -h
# Your home directory size
du -sh ~/
# Find what's taking the most space in your home directory
du -sh ~/* | sort -rh | head -20
# If you used &
jobs
# Your Python processes specifically
ps -u $USER -f | grep python
# Your tmux sessions
tmux ls
# Your screen sessions
screen -ls
This is the most important section of this guide. The SDS server is a shared resource with no job scheduler, which means a single user’s runaway job can bring the entire server down for everyone.
nice your jobsThe nice command tells the operating system to give your
job lower scheduling priority. This means if someone is running an
interactive session, your batch job won’t freeze them out.
# Good
nice python my_simulation.py
# Bad — runs at normal priority, hogs resources
python my_simulation.py
For even lower priority:
nice -n 10 python my_simulation.py
This is the single most common way the server gets
crashed. Many libraries in Python and R will, by default, use
every available CPU core. On this machine os.cpu_count()
and detectCores() both return 255, so a
single job left at its defaults can spawn 255 workers and consume the
entire server.
NumPy, SciPy, and other libraries using BLAS/OpenMP:
These libraries use multi-threaded linear algebra under the hood. Limit them with environment variables:
export OMP_NUM_THREADS=8
export MKL_NUM_THREADS=8
export OPENBLAS_NUM_THREADS=8
Put these in your ~/.bashrc so they apply every time you
log in:
echo 'export OMP_NUM_THREADS=8' >> ~/.bashrc
echo 'export MKL_NUM_THREADS=8' >> ~/.bashrc
echo 'export OPENBLAS_NUM_THREADS=8' >> ~/.bashrc
source ~/.bashrc
PyTorch DataLoader:
# Bad — os.cpu_count() returns 255 on this machine, so this spawns 255 workers
DataLoader(dataset, num_workers=os.cpu_count())
# Good — use a reasonable number
DataLoader(dataset, num_workers=4)
multiprocessing and
joblib:
# Bad
from multiprocessing import Pool
pool = Pool() # Defaults to ALL cores
# Good
pool = Pool(8) # Use 8 cores
# With joblib
from joblib import parallel_backend
with parallel_backend('loky', n_jobs=8):
results = Parallel()(delayed(func)(x) for x in data)
Pandas:
Some pandas operations are parallelized internally. Control this with
the same OMP_NUM_THREADS variable above.
parallel package:
# Bad — detectCores() returns 255 on this machine
library(parallel)
mclapply(data, func, mc.cores = detectCores())
# Good — leave cores for others
mclapply(data, func, mc.cores = 8)
foreach and
doParallel:
# Bad
library(doParallel)
registerDoParallel(cores = detectCores())
# Good
registerDoParallel(cores = 8)
BLAS threads (R uses these for matrix operations):
Add to your ~/.Renviron file:
OMP_NUM_THREADS=8
MKL_NUM_THREADS=8
OPENBLAS_NUM_THREADS=8
General rule of thumb: Use no more than 8-16 physical cores
per job unless you have confirmed that no one else is running
anything. Check with htop before deciding.
The server has roughly 881 GiB of RAM, shared across all users. Loading very large datasets entirely into memory, or running many parallel workers each with their own copy of data, can exhaust memory and take the machine down for everyone.
free -h and htopMemory is the one you can actually break. Oversubscribing the CPUs makes everyone slow, which is annoying but recoverable. Exhausting memory is worse: when the machine runs out, the Linux “OOM killer” starts terminating processes to recover — and it does not necessarily pick yours. Someone else’s three-day simulation can die because of your data frame. Before a large job, check
free -hand stay well within what is free.
Python example — memory-mapped files for large data:
import numpy as np
# Instead of loading everything into RAM:
data = np.load("huge_array.npy", mmap_mode='r')
R example — chunked reading:
library(data.table)
# Instead of read.csv (loads everything at once):
dt <- fread("large_file.csv", nrows = 100000) # Read first 100k rows
Storage on the server is limited and shared. Large files from one user can fill up the disk and cause problems for everyone, including the system itself.
conda env remove -n old_envCheck the whole disk and your own footprint:
df -h # usage of every filesystem — look at the one holding /home
du -sh ~/ # total size of your home directory
List each folder in your home directory, biggest first. Include
hidden folders like .cache and .conda — on
this server they are usually the real culprits:
# Top 10 folders, visible and hidden, largest first
du -sh ~/.[!.]* ~/* 2>/dev/null | sort -rh | head -10
The 2>/dev/null silences “permission denied” noise.
The biggest item is almost always ~/miniforge3 (your Conda
install), followed by dataset and checkpoint folders.
A miniforge install easily grows to tens of GB. First see where the space goes:
du -sh ~/miniforge3/envs/* | sort -rh # size of each environment
conda env list # environments you still have
Then reclaim space in two safe steps:
# 1. Clear Conda's download + package cache
conda clean --all
# 2. Delete whole environments you no longer use (the biggest win)
conda env remove -n old_env
Is
conda cleansafe? Yes.conda clean --allasks separately about tarballs, the index cache, and packages. All three act only on the package cache in~/miniforge3/pkgs— downloaded archives and unpacked entries that no environment is currently using. It does not remove packages from your active environments, so answeringyto all three is safe. The only cost is that reinstalling a cleaned package later re-downloads it. Deleting an entire unused environment withconda env removefrees far more than cache cleaning ever will.
nvidia-smi before starting a GPU jobCUDA_VISIBLE_DEVICES to use only the GPUs you need
(usually 1)These numbers are community etiquette, not enforced limits — the server has no scheduler to stop you, and they are not official policy. When in doubt (e.g., you genuinely need more for a deadline), check with the server administrator or coordinate with other users first.
| Resource | Recommended Limit |
|---|---|
| CPU cores per job | 8-16 physical cores (check htop first) |
| Total CPU cores per user | No more than ~32 unless the server is idle |
| RAM per job | Check free -h; stay under ~150 GB |
| GPUs per user | 1-2 (check nvidia-smi first) |
PyTorch num_workers |
4-8 — never os.cpu_count(), which returns 255 here |
| Home directory size | Clean up regularly; no permanent bulk storage |
CNS-SRV-SDS01.austin.utexas.eduping CNS-SRV-SDS01.austin.utexas.edussh-add -ltmux or screen (see Running Jobs)tmux new -s myjob before running your job, then
Ctrl+b, d to detachnohup is not reliable on this
server — jobs started with it may still die at logoutconda activate myprojectinstalled.packages()nvidia-smi shows my GPU process but the job
crashedThe process might be a zombie. Kill it:
# Find your GPU processes
nvidia-smi
# Kill by PID
kill PID
# If it won't die
kill -9 PID
Someone (maybe you) might be using too many cores. Check:
htop
If the CPU meters are all pegged and the load average is well above 128, find the top consumers and who owns them:
ps aux --sort=-%cpu | head -15
See Figuring out who is using resources for more ways to break down usage by user. Then talk to the person (or email the admin) — don’t kill anyone else’s processes.
nvidia-smi shows the PIDs of GPU processes but not their
owners. Map the PIDs to usernames and commands:
ps -up $(nvidia-smi --query-compute-apps=pid --format=csv,noheader)
Then coordinate with that person directly, or wait for their job to finish — never kill another user’s process.
The disk is full. Check your usage:
du -sh ~/
Clean up what you can. If you’re not the cause, contact the admin.
Reduce your batch size, or use a different (less occupied) GPU:
CUDA_VISIBLE_DEVICES=2 python train.py --batch_size 16
For questions or issues, contact the SDS server administrator. For suggested edits to this guide, reach out to Ritwik Vashistha.