Working with Apptainer containers
How to run a container (review)
As a reminder, we continue working in our ~/tmp directory inside an interactive job with the Apptainer module loaded:
cd ~/tmp
module load apptainer
salloc --time=2:00:0 --mem-per-cpu=3600 # only from the login nodeIf you have not done so already, let’s pull the latest Ubuntu container from Docker:
apptainer pull ubuntu.sif docker://ubuntuWe already saw some of these commands:
apptainer shell ubuntu.siflaunches the container and opens an interactive shell inside itapptainer exec ubuntu.sif <command>launches the container, runs a single command inside it and then exitsapptainer run ubuntu.siflaunches the container, executes the default runscript and then exits
Apptainer matches users between the container and the host. For example, if you run a container that needs to be root, you also need to be root outside the container.
1. Running a single command
apptainer exec ubuntu.sif ls /
apptainer exec ubuntu.sif ls /; whoami
apptainer exec ubuntu.sif bash -c "ls /; whoami" # probably a safer way
apptainer exec ubuntu.sif cat /etc/os-release2. Running a default script
We’ve already done this! If there is no default script, Apptainer will give you the shell to type in your commands.
3. Starting a shell
We’ve already done this!
$ apptainer shell ubuntu.sif
Apptainer> whoami # same username as in the host system
Apptainer> groups # tries to match groups as on the host systemAt startup Apptainer simply copied the relevant user and group lines from the host system to files /etc/passwd and /etc/group inside the container. Why do this? The container must ensure that you cannot modify anything on the host system that you should not have permission to, i.e. you are restricted to the same user permissions within the container as you are on the host system.
Mounting external directories and copying variables
By default, Apptainer containers are read-only, so you cannot write into its directories. However, from inside the container you can organize read-write access to your directories on the host filesystem. The command
apptainer shell -B /home,/project,/scratch ubuntu.sifwill bind-mount /home,/project,/scratch inside the container so that these directories can be accessed for both read and write, subject to your account’s permissions, and then will run a shell. Inside the container:
pwd # most likely your working directory
echo $USER
ls /home
ls /scratch
ls /projectIf you prefer, you can pass the same bind information via an environment variable APPTAINER_BIND. In fact, by default your APPTAINER_BIND is set to /project,/scratch, so these two directories (along with /home/$USER) will be mounted every time. If you want to stop mounting /project and /scratch, unset the variable:
unset APPTAINER_BIND
apptainer shell ubuntu.sifYou can mount host directories to specific paths inside the container, e.g.
apptainer shell -B /project/def-sponsor00/$USER:/myproject,/home/$USER/scratch:/myscratch ubuntu.sif
Apptainer> ls /myproject
Apptainer> ls /myscratchNote that by default Apptainer typically mounts some of the host’s directories (think /home/$USER). The flag -C will hide the host’s filesystems and environment variables, but then you need to explicitly bind-mount the needed paths (to store results), e.g.
apptainer shell -C -B /scratch ubuntu.sif # from the host see only /scratch
Apptainer> ls /home/$USER # still there, but does not contain host's files and directoriesYou can disable specific mounts, e.g. the following will start the container without mounting your home directory, but it’ll mount the current directory:
apptainer shell --no-mount home ubuntu.sifAlternatively, you can disable mounting /home with the --no-home flag, which is equivalent to --no-mount home. And you can disable multiple mounts with something like --no-mount tmp,sys,dev.
In general, without -C, Apptainer inherits all environment variables and default bind-mounted filesystems. You can add the -e flag to remove only the host’s environment variables from your container but keep the default bind-mounted filesystems, to start in a cleaner environment:
apptainer shell ubuntu.sif
Apptainer> echo $USER $PYTHONPATH # defined from the host
Apptainer> ls /home/user01 # shows my $HOME content on the host
apptainer shell -C ubuntu.sif
Apptainer> echo $USER $PYTHONPATH # not defined
Apptainer> ls /home/user01 # nothing
apptainer shell -e ubuntu.sif
Apptainer> echo $USER $PYTHONPATH # not defined
Apptainer> ls /home/user01 # shows my $HOME content on the hostOn the other hand, you can pass variables to your container by prefixing their names:
APPTAINERENV_HI="hello" APPTAINERENV_NAME="alex" apptainer shell ubuntu.sif
Apptainer> echo $HI # should print hello
Apptainer> echo $NAME # should print alexFinally, we already mentioned APPTAINER_BIND: you don’t have to pass the same bind (-B) flags every time – instead you can put them into a variable (that can be stored in your ~/.bashrc file):
export APPTAINER_BIND="/home,/project/def-sponsor00/${USER}:/project,/scratch/${USER}:/scratch"
apptainer shell ubuntu.sifYou can have more granular control (e.g. specifying read only) with the --mount flag – for details see the official Bind Paths and Mounts documentation.
- Your current directory and home directory are usually available by default in a container.
- You have the same username and permissions in a container as on the host system.
- Use
-Bto mount host’s directories inside the container. - Use
-Cto hide both host’s filesystems and environment variables, perhaps while mounting only few specific directories. - Use
-eto hide only the host’s environment variables.
Overlays
In the container world, an overlay image is a file formatted as a filesystem. To the host filesystem it is a single file. When you mount it into a container, the container will see a filesystem with many files.
An overlay mounted into an immutable SIF image lets you store files without rebuilding the image. For example, you can store your computation results, or compile/install software into an overlay.
An overlay can be:
- a standalone writable ext3 filesystem image (most useful),
- a sandbox directory, or
- a writable ext3 image embedded into the SIF file.
If you write millions of files, do not store them on a cluster filesystem – instead, use an Apptainer overlay file for that. Everything inside the overlay will appear as a single file to the cluster filesystem which will help to reduce I/O issues on Lustre.
cd ~/tmp
apptainer pull ubuntu.sif docker://ubuntu:latest
module load apptainer
salloc --time=0:30:0 --mem-per-cpu=3600
apptainer overlay create --size 512 small.img # create a 0.5GB overlay image file
apptainer shell --overlay small.img ubuntu.sifInside the container any newly-created top-level directory will go into the overlay filesystem:
Apptainer> df -kh # the overlay should be mounted inside the container
Apptainer> mkdir -p /data # by default this will go into the overlay image
Apptainer> cd /data
Apptainer> df -kh . # using overlay; check for available space
Apptainer> for num in $(seq -w 00 19); do
echo $num
# generate a binary file (1-33)MB in size
dd if=/dev/urandom of=test"$num" bs=1024 count=$(( RANDOM + 1024 ))
done
Apptainer> df -kh . # should take ~300-400 MBIf you exit the container and then mount the overlay again, your files will be there:
apptainer shell --overlay small.img ubuntu.sif
Apptainer> ls /data # here is your dataYou can also create a new overlay image with a directory inside with something like:
apptainer overlay create --create-dir /data --size 512 overlay.img # create an overlay with a directoryIf you want to mount the overlay in the read-only mode:
apptainer shell --overlay small.img:ro ubuntu.sif
Apptainer> touch /data/test.txt # error: read-only file systemInto the same container at the same time, you can mount many read-only overlays, but only one writable overlay.
To see the help page on overlays (these two commands are equivalent):
apptainer help overlay create
apptainer overlay create --helpSparse overlay images
Sparse images use disk more efficiently when blocks allocated to them are mostly empty. As you add more data to a sparse image, it can grow (but not shrink!) in size. Let’s create a sparse overlay image:
apptainer overlay create --size 512 --sparse sparse.img
ls -l sparse.img # its apparent size is 512MB
du -h --apparent-size sparse.img # same
du -h sparse.img # its actual size is much smaller (316K)Let’s mount it and fill with some data, e.g. a few files:
apptainer shell --overlay sparse.img ubuntu.sif
Apptainer> mkdir -p /data && cd /data
Apptainer> for num in $(seq -w 0 4); do
echo $num
# generate a binary file (1-33)MB in size
dd if=/dev/urandom of=test"$num" bs=1024 count=$(( RANDOM + 1024 ))
done
Apptainer> df -kh . # should take ~75-100 MB, pay attention to "Used"
du -h sparse.img # shows actual usage
Be careful with sparse images: not all tools (e.g. backup/restore, scp, sftp, gunzip) recognize sparsefiles ⇒ this can potentially lead to data loss and other bad things …
Example: installing Conda into an overlay
Installing native Anaconda on HPC clusters is a bad idea for a number of reasons. Instead of Conda, we recommend using virtualenv together with our pre-compiled Python wheels to install Python packages into your own virtual environments.
One of the reasons we do not recommend Conda is that it creates a large number of files in your directories. You can alleviate this problem by hiding Conda files inside an overlay image:
- takes a couple of minutes, results in 32k+ files that are hidden from the host
- no need for root, as you don’t modify the container image
- still might not be the most efficient use of resources (non-optimized binaries)
Here is one way you could install Conda into an overlay image:
cd ~/tmp
apptainer pull ubuntu.sif docker://ubuntu:latest
apptainer overlay create --size 1200 conda.img # create a 1200M overlay image
wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh -O miniconda.sh
apptainer shell --overlay conda.img -B /home ubuntu.sif
Apptainer> mkdir /conda && cd /conda
Apptainer> df -kh .
Apptainer> bash /home/${USER}/tmp/miniconda.sh
agree to the license
use /conda/miniconda3 for the installation path
no to initialize Miniconda3
Apptainer> find /conda/miniconda3/ -type f | wc -l # 32,621 files
Apptainer> df -kh . # uses ~840M once finished, but more was used during installationThese 32k+ files appear as a single file to the Lustre metadata server (which is great!).
apptainer shell ubuntu.sif
Apptainer> ls /conda # no such file or directory
apptainer shell --overlay conda.img ubuntu.sif
Apptainer> source /conda/miniconda3/bin/activate
(base) Apptainer> type python # /conda/miniconda3/bin/python
(base) Apptainer> python # worksIf you want to install large Python packages, you probably want to resize the image:
e2fsck -f conda.img # check your overlay's filesystem first (required step)
resize2fs -p conda.img 2G # resize your overlay
ls -l conda.img # should be 2GBNext mount the resized overlay into the container, and make sure to pass the -C flag to force writing config files locally (and not to the host):
apptainer shell -C --overlay conda.img ubuntu.sif # /home/$USER won't be available
Apptainer> cd /conda/miniconda3
Apptainer> source bin/activate
(base) Apptainer> conda install numpy
(base) Apptainer> df -kh . # so far used 1.8G out of 2.0G
(base) Apptainer> python
>>> import numpy as np
>>> np.piContainer best practices
- Should we put everything into one container? A monolithic container (one SIF file) is easier to build when starting from scratch, but would need to be rebuilt to update one of its components. A modular container makes it easier to swap and upgrade individual components as needed. Consider this:
apptainer run --overlay {layer1,layer2,layer3}.img container.sif - Use overlays to hide many files from the host’s filesystem.
- Use trusted images, e.g. you should probably pick
python:3.12overfoobar/python:3.12. - When building images, only keep what’s needed, e.g. don’t forget to remove source code and other dependencies after compilation, clean package manager caches after installing software, start with a minimal base image appropriate for your task.
- For GPU/AI workloads, try building on the base CUDA or HIP/ROCM image, or pull an existing application container from NVIDIA NGC (see next section).