Installation¶
ADToolbox runs as a locally installed Python package or from a container image. Anything
that shells out to an external bioinformatics tool works either way: put the tool on your
PATH, or let ADToolbox call it inside Docker or Apptainer/Singularity. Pipeline steps can
additionally be submitted to Slurm, which is how the toolbox scales to HPC clusters.
Use a virtual environment
ADToolbox pins several scientific dependencies. Install it into a dedicated environment rather than your system Python:
Requirements: Python 3.11 or newer.
Install the package¶
The environment file installs ADToolbox and every external tool both the
amplicon and shotgun routes need, so nothing else has to be on your PATH:
curl -O https://raw.githubusercontent.com/chan-csu/ADToolbox/main/environment.yml
conda env create -f environment.yml # or: mamba env create -f environment.yml
conda activate adtoolbox
This is the least-fuss way to get a fully working pipeline — it pulls fastp, Cutadapt, DADA2, VSEARCH, MMseqs2, SRA Toolkit, and the NCBI Datasets CLI alongside the Python package, so you can skip the External tools section entirely.
A conda install -c bioconda adtoolbox one-liner is on the way
The Bioconda recipe
is under review. Once it merges, conda install -c bioconda adtoolbox
will install the package directly; until then, use the environment file
above.
Installs the Python package only. The metagenomics pipeline additionally
needs the external tools on your PATH, or a container.
Verify the install:
Optional extras¶
The base install covers databases, the metagenomics pipeline, and ADM simulation. Features with heavier dependencies are opt-in:
| Extra | Install | Enables |
|---|---|---|
dashboard |
pip install "adtoolbox[dashboard]" |
Interactive Dash visualization, including Escher maps (--report dash). |
blackbox |
pip install "adtoolbox[blackbox]" |
BlackBoxOptimizer via OpenBox. |
genetic |
pip install "adtoolbox[genetic]" |
GeneticOptimizer via PyGAD. |
surrogate |
pip install "adtoolbox[surrogate]" |
SurrogateOptimizer via PyTorch. |
optimize |
pip install "adtoolbox[optimize]" |
All three optimizer backends at once. |
Combine them as needed:
External tools¶
These are only required for the metagenomics pipeline. If you plan to use containers, skip
this section — the ADToolbox image already contains all of them. The conda / mamba
install above also pulls every one of them.
| Tool | Used for | Assay |
|---|---|---|
| fastp | Adapter trimming and quality filtering. | amplicon + shotgun |
| MMseqs2 | Protein alignment against the ADToolbox enzyme database. | amplicon + shotgun |
| SRA Toolkit | prefetch and fasterq-dump for downloading reads from SRA. |
both (SRA input) |
| Cutadapt | IUPAC-aware PCR primer removal. | amplicon only |
| DADA2 and R | Error learning, ASV inference, paired-read merging, and chimera removal. | amplicon only |
| VSEARCH | Mapping representative ASVs to GTDB; also the legacy feature backend. | amplicon only |
| NCBI Datasets | Batched assembly downloads. | amplicon (genomes) |
Shotgun-only setups are lighter
The shotgun route needs just fastp and MMseqs2 (plus SRA Toolkit for
--input-type sra). If you never run the amplicon route, you can skip
Cutadapt, DADA2/R, VSEARCH, and the NCBI Datasets CLI.
The quickest local install is conda/mamba:
mamba install -c conda-forge -c bioconda \
fastp cutadapt r-base r-digest bioconductor-dada2 \
vsearch mmseqs2 ncbi-datasets-cli sra-tools
This installs standalone DADA2 rather than QIIME2. Set container = "None" in the execution profile so Slurm jobs use these commands from the active environment.
Containers¶
The published image bundles ADToolbox with every external tool.
You do not have to run ADToolbox itself inside the container. A local install can call the
container for individual pipeline steps with --container docker or
--container singularity, which is usually the more convenient arrangement:
adtoolbox metagenomics align-genome \
--name my_genome \
--input-file ./genomes/my_genome.fna \
--output-dir ./alignment \
--protein-db ./database/Protein_DB.fasta \
--container docker
HPC and Slurm¶
For cluster runs, describe each pipeline step in a TOML execution profile that selects the
backend, container, and resource request. ADToolbox then generates and submits the Slurm
scripts, waits for them with sbatch --wait, and can retry failed jobs.
backend = "local"
container = "None"
[steps.align_short_reads]
backend = "slurm"
cpus = 24
memory = "150G"
time = "12:00:00"
A complete reference profile ships at reference_data/metagenomics_pipeline.toml. See
Execution profiles for every option.
Run without installing¶
Binder launches the example notebooks in a browser with no local setup. Escher map visualization is not available there.
Next steps¶
Head to the Quickstart to download the reference databases and run your first simulation.