Configure paths for your machine
Use this guide before running the full-data workflows or manuscript figures.
The public toy tutorial does not require the external study datasets or this
configuration. Run the commands below in Bash, from the repository root.
Replace every /srv/... example with a real absolute path on your machine.
scripts/config.sh currently contains the authors' data and software roots.
You can edit its defaults or override them with exported environment variables.
Configuration selects files; it does not download data, create missing analysis
inputs, install software, or convert genome builds. See the pipeline guide
for data sources, access restrictions and column formats.
1. Choose where inputs and results live
| Setting | What it points to | Example |
|---|---|---|
TRACECB_DATA_ROOT |
Root containing external datasets and references. | /srv/tracecb/input |
TRACECB_SOFTWARE_ROOT |
Parent of external tools; defaults expect plink, plink2, and s-ldxr-master/s-ldxr.py below it. |
/srv/tracecb/software |
TRACECB_OUTPUT_ROOT |
Writable root for pipeline outputs. The population/tissue subdirectory is added automatically. | /srv/tracecb/runs/run01 |
TRACECB_STUDY_ROOT |
Root containing study results to read, including population/tissue subdirectories. | /srv/tracecb/runs/run01 |
TRACECB_STUDY_DIR |
One population/tissue directory containing QTD* study directories. |
/srv/tracecb/runs/run01/EAS_eQTLGen |
TRACECB_FIGURE_DIR |
Writable output directory for figures in the selected configuration. | /srv/tracecb/runs/run01/figures/EAS_eQTLGen |
TRACECB_LOG_DIR |
Pipeline logs; default is TRACECB_OUTPUT_ROOT/logs. |
/srv/tracecb/runs/run01/logs |
TRACECB_REPO_ROOT |
Repository checkout; normally detected automatically. | /srv/code/traceCB |
Pipeline output and figure input are separate settings. GMM writes to
TRACECB_OUTPUT_ROOT/<population>_<tissue>/QTD*/GMM/. Figures read from
TRACECB_STUDY_DIR, whose default is under TRACECB_DATA_ROOT/traceCB, not under
the output root. To plot a new run, explicitly set TRACECB_STUDY_ROOT to its
output root before sourcing the configuration. Do not append EAS_eQTLGen
to TRACECB_OUTPUT_ROOT; otherwise the wrapper adds that component twice.
Use a new output directory for a new analysis version. Existing files are not automatically protected from overwriting, and stale outputs can otherwise be mixed with newly generated results.
2. Save a machine-specific configuration
Keep local paths outside the checkout, for example in ../tracecb-local.sh.
This sample assumes the directory layout in section 3:
# Contents of ../tracecb-local.sh; replace paths before sourcing.
export TRACECB_DATA_ROOT="/srv/tracecb/input"
export TRACECB_SOFTWARE_ROOT="/srv/tracecb/software"
export TRACECB_OUTPUT_ROOT="/srv/tracecb/runs/run01"
export TRACECB_STUDY_ROOT="$TRACECB_OUTPUT_ROOT"
export TARGET_POPULATION="EAS" # EAS or AFR
export TISSUE_SOURCE="eQTLGen" # eQTLGen or GTEx; case matters
export PYTHON_ENV="py312" # existing Conda environment name
export R_ENV="r4" # needed for the optional coloc wrapper
# These overrides are optional if the software root has the default layout.
export PLINK_BIN="/srv/tracecb/software/plink"
export PLINK2="/srv/tracecb/software/plink2"
export SLDXR_DIR="/srv/tracecb/software/s-ldxr-master"
export TRACECB_GTEX_GENE_ANNOTATION="/srv/tracecb/input/GTEx/gencode.v26.GRCh38.genes.gtf"
export CELL_TYPE_PROPORTION_FILE="/srv/tracecb/input/GTEx/celltype_proportion.csv"
export MAX_JOBS=8
Apply it in a shell that has not already sourced a different traceCB configuration:
source ../tracecb-local.sh
source scripts/config.sh
source scripts/config.sh exports figure paths, sets PYTHONPATH and the
headless Matplotlib backend, and creates output/log directories. It does not
activate a Conda environment. The pipeline shell wrappers activate PYTHON_ENV
or R_ENV themselves; direct Python/R plotting commands use the active environment.
Do not run bash scripts/config.sh: it cannot configure the calling shell and
the script deliberately rejects that invocation.
Environment values take precedence over defaults written as ${VARIABLE:-default}.
Use export so overrides reach wrappers started with bash scripts/....
Do not just assign DATA_ROOT or OUTPUT_ROOT; the corresponding user overrides
are TRACECB_DATA_ROOT and TRACECB_OUTPUT_ROOT.
When changing population, tissue or roots, start from a clean environment and
source the appropriate local settings again. A child bash launched from an
already configured shell inherits its exports and is not a clean reset.
Re-sourcing only TARGET_POPULATION=AFR does not recompute existing exported
TRACECB_STUDY_DIR, TRACECB_FIGURE_DIR, coloc paths or other derived paths.
If reusing a configured shell, explicitly update or unset every affected override.
3. Supply the inputs for the stages you will run
Harmonization, LD scores and GMM
The defaults expect the following layout. Only the selected target population
and tissue source are needed. chr1 examples repeat through chr22.
input/
├── BBJ_eQTL/by_celltype_chr/Monocytes/chr1.csv
├── popcell/AFB_NS/MONO__NS/eQTL_AFB_chr1_assoc.txt.gz
├── eQTLCatalogue/by_celltype_chr/QTD000021/chr1.csv
├── eQTLGen/chr1.tsv.gz
├── GTEx/
│ ├── GTEx_Whole_Blood_by_chr/chr1.csv
│ ├── celltype_proportion.csv
│ └── gencode.v26.GRCh38.genes.gtf
└── 1000G/
├── 1000G_EAS/1000G.EAS.QC.maf.1.bed
├── 1000G_EAS/1000G.EAS.QC.maf.1.bim
├── 1000G_EAS/1000G.EAS.QC.maf.1.fam
├── 1000G_EUR/1000G.EUR.QC.maf.1.{bed,bim,fam}
└── 1000G_AFR/1000G.AFR.QC.maf.1.{bed,bim,fam}
{bed,bim,fam} denotes three files, not a literal filename. Add all required
cell types and studies, following STUDY_IDS and CELL_TYPES in the config.
| Override | Required contents / role |
|---|---|
BBJ_DIR |
Formatted EAS target inputs: <cell-type>/chr<chr>.csv, e.g. Monocytes/chr22.csv. Defaults to DATA_ROOT/BBJ_eQTL/by_celltype_chr. |
AFR_DIR |
Formatted AFR target inputs: <label>__NS/eQTL_AFB_chr<chr>_assoc.txt.gz. Labels are MONO, B, NK, T.CD4, T.CD8. |
EQTL_CATALOGUE_DIR |
Formatted auxiliary EUR inputs: <QTD study>/chr<chr>.csv. |
EQTLGEN_DIR |
eQTLGen chr<chr>.tsv.gz, tab-delimited and retaining the original headers expected by load_eQTLGen(). A whole-genome download is not directly interchangeable with these files. |
GTEX_SOURCE_DIR |
Raw GTEx download directory; default parent for GTEx files below. |
GTEX_DIR |
Formatted GTEx tissue inputs: chr<chr>.csv. |
CELL_TYPE_PROPORTION_FILE |
CSV named celltype_proportion.csv, with Cell_type,Proportion columns. Values are percentages: 20 means 20%, not 0.2. Use cell-type names matching the config. The preparation wrapper copies the basename unchanged, and GMM expects this exact basename. |
LD_REFERENCE_DIR |
Parent of 1000G_EAS, 1000G_AFR and, by default, 1000G_EUR. |
AUX_LD_DIR |
Directory containing the EUR reference BED/BIM/FAM triplets; can override the default LD_REFERENCE_DIR/1000G_EUR. |
The loader currently chooses input formats using directory names: the target
path must contain bbj or af, and the tissue path must contain gtex or
eqtlgen (case-insensitive). Keep these identifiers when choosing custom paths.
TARGET_EQTL_DIR, TISSUE_DIR, AUX_EQTL_DIR, TARGET_LD_DIR and OUTPUT_DIR
are derived assignments that the config overwrites. Set their parent settings
from the table instead. In particular, exporting TARGET_LD_DIR alone does not
override the selected LD_REFERENCE_DIR/1000G_<population> directory. For a
different target layout, arrange the selected subdirectory accordingly or use
the underlying Python tools' explicit path arguments.
Put LD triplets directly in the configured population directory. The loader
expects 1000G.*.QC.maf.<chr>.bim; there should be one matching reference per
chromosome. A download named 1000G.EUR.QC.<chr> does not match that pattern.
Confirm ancestry, genome build, QC and allele conventions before adapting a
reference layout; changing a filename does not perform QC or liftover.
harmonize_inputs.py matches rsIDs and retains tissue coordinates without liftover.
Coordinate-based coloc/enrichment inputs must use matching builds.
Study IDs, cell types, sample sizes and the chromosome list are Bash arrays or
unconditional assignments in config.sh, not environment-variable overrides.
For your own cohorts, edit those coordinated lists and any manuscript figure
metadata that assumes the paper's ten studies. Exporting CHROMOSOMES=22 will
not restrict the wrappers. For one prepared study/chromosome, use the CLI:
python -m traceCB.run_gmm \
--study QTD000021 --cell-type Monocytes --chromosome 22 \
--data-dir "$TRACECB_OUTPUT_ROOT/${TARGET_POPULATION}_${TISSUE_SOURCE}"
This command requires already harmonized inputs, gene LD scores and proportions under that study directory; it does not prepare them.
Raw-data preprocessing settings
These settings belong to the preprocessing scripts. Their outputs feed the formatted-input paths above; setting one side does not always update the other.
| Script | Raw input settings | Output setting and next-stage input |
|---|---|---|
preprocess_bbj.sh <cell-type> |
BBJ_SOURCE_DIR/eQTL_<cell-type>.tar.gz. |
BBJ_OUTPUT_DIR; set BBJ_DIR to the same directory when overriding. |
preprocess_eqtl_catalogue.sh |
EQTL_CATALOGUE_SOURCE_DIR containing Catalogue study files. |
EQTL_CATALOGUE_DIR. |
preprocess_gtex.sh |
GTEX_EQTL_FILE and GTEX_LOOKUP_FILE (both gzip files). TRACECB_GTEX_LOOKUP also supplies the lookup default; keep overrides consistent. |
GTEX_OUTPUT_DIR; set GTEX_DIR to the same directory when overriding. |
preprocess_1000g.sh |
REFERENCE_DIR containing all_hg38.pgen, all_hg38.pvar, all_hg38.psam, or the supported compressed source files. |
Population subdirectories under REFERENCE_DIR; set LD_REFERENCE_DIR to this parent for subsequent stages. Requires PLINK1 and PLINK2. |
The exact default GTEx lookup filename is
GTEx_Analysis_2017-06-05_v8_WholeGenomeSeq_838Indiv_Analysis_Freeze.lookup_table2017.48.22.txt.gz.
The default association filename is
GTEx_Analysis_v8_QTLs-GTEx_Analysis_v8_eQTL_all_associations-Whole_Blood.allpairs.txt.gz.
Override their variables if your downloaded filenames differ.
The repository does not currently provide a dedicated eQTLGen chromosome-splitting
entry point; prepare the specified chromosome files before running harmonization.
External software
| Setting | Required value |
|---|---|
PLINK_BIN |
PLINK 1.9 executable path, not its containing directory. A command name on PATH also works. |
PLINK2 |
PLINK 2 executable, used by 1000G preprocessing. |
PLINK1 |
Optional PLINK 1.9 override specific to 1000G preprocessing; otherwise uses PLINK_BIN. |
SLDXR_DIR |
Directory containing s-ldxr.py; its Python dependencies must be installed in the environment running LD scores. |
PYTHON_ENV / R_ENV |
Existing Conda environment names. The Python environment.yml does not create the R environment. |
PYTHON_BIN |
Optional Python executable used by the main workflow wrappers after environment activation. Prefer the activated environment's python. |
MAX_JOBS |
Positive integer for the wrappers that implement this limit: GTEx preprocessing, LD scores, GMM and coloc. Input preparation instead launches one process per configured study. |
See the pipeline guide and the figure README for dependencies. Record actual PLINK/S-LDXR/R versions and commits with your run; paths alone do not freeze software versions.
4. Check the resolved paths before expensive jobs
After sourcing both configuration files, inspect what the wrappers will use:
for name in TARGET_POPULATION TISSUE_SOURCE TARGET_EQTL_DIR AUX_EQTL_DIR \
TISSUE_DIR CELL_TYPE_PROPORTION_FILE TARGET_LD_DIR AUX_LD_DIR OUTPUT_DIR \
TRACECB_STUDY_DIR TRACECB_FIGURE_DIR PLINK_BIN SLDXR_DIR; do
printf '%s=%s\n' "$name" "${!name}"
done
# Main prerequisites before input preparation / LD scoring:
(
set -e
test -d "$TARGET_EQTL_DIR"
test -d "$AUX_EQTL_DIR"
test -d "$TISSUE_DIR"
test -s "$CELL_TYPE_PROPORTION_FILE"
command -v "$PLINK_BIN"
test -s "$SLDXR_DIR/s-ldxr.py"
# Require exactly one correctly named reference and its complete triplet.
shopt -s nullglob
for ref_dir in "$TARGET_LD_DIR" "$AUX_LD_DIR"; do
for chr in "${CHROMOSOMES[@]}"; do
bims=("$ref_dir"/1000G.*.QC.maf."$chr".bim)
if (( ${#bims[@]} != 1 )); then
printf 'Expected one BIM for chr%s in %s; found %s\n' \
"$chr" "$ref_dir" "${#bims[@]}" >&2
exit 1
fi
for ext in bed bim fam; do
test -s "${bims[0]%.bim}.$ext" || exit 1
done
done
done
)
These checks cover paths and filenames, not scientific validity or full CSV schemas. Resolve failures before continuing. A pipeline directory may exist because the config created it even when no analyses have run.
Then run the needed stages in order:
bash scripts/prepare_inputs.sh
bash scripts/run_ld_scores.sh
bash scripts/run_gmm.sh
Before making full-study figures, check completeness explicitly. The plotting loader can otherwise read a subset of chromosomes without treating it as an error:
(
missing=0
for study in "${STUDY_IDS[@]}"; do
for chr in "${CHROMOSOMES[@]}"; do
file="$TRACECB_STUDY_DIR/$study/GMM/chr$chr/summary.csv"
if [[ ! -s "$file" ]]; then
printf 'Missing or empty: %s\n' "$file" >&2
missing=1
fi
done
done
exit "$missing"
)
This only checks presence and size. Also inspect logs, summary columns/rows and per-gene Parquet outputs. A header-only summary is not evidence of a successful full analysis. Label deliberately partial analyses as partial.
5. Configure manuscript figures
For case-study plots of a newly generated run, the local example already points the study input at the output root. Install the figure dependencies and activate the intended Python environment before running:
conda activate py312
pip install -e '.[figures]'
source ../tracecb-local.sh
source scripts/config.sh
python -m figures.case_study
python -m figures.combine_case_study
For existing results, use these overrides instead, in a clean shell:
export TRACECB_DATA_ROOT="/srv/tracecb/input"
export TRACECB_OUTPUT_ROOT="/srv/tracecb/figure-runs/run01"
export TRACECB_STUDY_ROOT="/srv/tracecb/archive"
export TARGET_POPULATION="EAS"
export TISSUE_SOURCE="eQTLGen"
export TRACECB_GTEX_GENE_ANNOTATION="/srv/tracecb/input/GTEx/gencode.v26.GRCh38.genes.gtf"
source scripts/config.sh
# Reads /srv/tracecb/archive/EAS_eQTLGen/QTD*/GMM/...
# Writes /srv/tracecb/figure-runs/run01/figures/EAS_eQTLGen/...
TRACECB_GTEX_GENE_ANNOTATION is a file, not a directory; the current reader
expects an uncompressed GTF. It is also required for gene annotation when the
tissue source is eQTLGen. TRACECB_FIGURE_METADATA defaults to the repository's
src/figures/metadata.json and normally should not be changed.
Only configure the following extra inputs for figures that consume them:
| Setting | Required input / use |
|---|---|
TRACECB_AFR_STUDY_DIR |
AFR cohort root with QTD*/GMM/chr*/summary.csv; defaults to STUDY_ROOT/AFR_<tissue>. |
TRACECB_ONEK1K_FILE |
Processed onek1k_esnp.csv, not an arbitrary raw OneK1K download. |
TRACECB_OASIS_DIR |
OASIS directory with all five <cell>_PC15_MAF0.05_Cell.10_top_assoc_chr1_23.txt.gz files; <cell> is Mono, CD4T, CD8T, B, NK. |
TRACECB_GTEX_LOOKUP |
Full GTEx variant lookup gzip; used for SNP matching in replication plots. |
TRACECB_CIMA_DIR |
Parent of the CIMA files below; override the individual files if their layout differs. |
TRACECB_CIMA_LEAD_EQTL |
CIMA_Lead_cis-xQTL.csv. |
TRACECB_CIMA_CELL_TYPES |
CIMA_Cell_Type_Level_and_Marker.xlsx; the openpyxl Excel reader is included in the figures extra. |
TRACECB_REPLICATION_EGENES / TRACECB_REPLICATION_ESNPS |
Processed hum0343_eGene.csv / hum0343_eSNP.csv. |
TRACECB_CELL_PROPORTIONS |
Individual-level ind_celltype_proportion.csv for the proportion figure; distinct from GMM's mean celltype_proportion.csv. |
TRACECB_TIMING_FILE |
summary_timing.csv in the schema consumed by runtime_comparison.py; a raw GMM text log is not a substitute. |
TRACECB_COLOC_INPUT_DIR |
Prepared GWAS/locus directory containing bcx/ and bbj/. Keep COLOC_INPUT_DIR consistent if also set. |
TRACECB_COLOC_DIR |
Generated *_coloc.csv result directory. Setting it selects figure input; the coloc shell wrapper writes to OUTPUT_DIR/coloc. |
TRACECB_COLOC_GENES |
Closest-gene BED file, e.g. bcx_mon.closest.protein_coding.bed. |
TRACECB_COLOC_REPLICATION |
Separately generated replication table, default COLOC_DIR/replication.csv; the main coloc wrapper does not create it. |
TRACECB_LOCUS_GWAS |
Prepared locus GWAS CSV, e.g. bcx_mon_GWAS.csv. |
TRACECB_LOCUS_EQTL_DIR |
Root of exported per-gene CSVs in the layout used by manhattan_locuszoom.R; GMM Parquet files alone are insufficient. |
TRACECB_LOCUS_TRACK_DIR |
ENCODE bigWig track directory, with the filenames expected by the locus scripts. |
When TRACECB_STUDY_ROOT points to a new run, unchanged external resources such
as onek1k_supp/, cell_type_proportion/ and prepared coloc/ inputs may still
be stored elsewhere. Override their individual paths; do not assume GMM creates
them in the new study root.
Optional output overrides are TRACECB_AFR_FIGURE_DIR,
TRACECB_ESNP_REPLICATION_DIR and TRACECB_COLOC_FIGURE_DIR. Their defaults are
afr/, esnp_replication/ and coloc/ under TRACECB_FIGURE_DIR. The full script
inventory and R packages are documented in src/figures/README.md.
6. Simulation and enrichment use additional settings
scripts/config.sh does not configure every simulation/enrichment path.
For the small-window simulations, use their own overrides explicitly:
SIM_DATA_DIR=/srv/tracecb/input/simulation \
OUT_DIR=/srv/tracecb/runs/run01/simulation \
IMG_DIR=/srv/tracecb/runs/run01/figures/simulation \
NREP=100 NSNP=2000 OMEGA_MODE=both \
bash src/simulation/run_all.sh all
SIM_DATA_DIR must contain EAS_n5000_chr22_loci29.npy and
EUR_n20000_chr22_loci29.npy, or set POP1_GENO and POP2_GENO to individual files.
The matrices are not distributed with the repository. Reproducing the exact
paper inputs additionally requires their data-generation/selection provenance.
Set CONDA_ENV for these wrappers; they do not use PYTHON_ENV.
Set SKIP_CONDA=1 only if the active environment is already suitable.
The all target does not include the whole-chromosome suite. Its
src/simulation/chr22/run.sh entry point also needs PLINK, SNP-gene and LD inputs;
SIM_DATA_DIR does not redirect its hard-coded defaults. Use the underlying
chr22/simulate.py --help path arguments for a custom input layout. Explicitly
set NREP and IMG_DIR: the shared setup currently supplies NREP=100, so the
later NREP=1 fallback in chr22/run.sh does not take effect. See
src/simulation/README.md for the separate grids and plotting commands.
S-LDSC enrichment uses STUDY_DIR, RESULT_DIR, GWAS_ROOT, LDSC_DIR and LD
reference prefixes, described in src/enrichment/README.md. LDSC_DIR is the
LDSC checkout, distinct from SLDXR_DIR. Its EUR reference stack uses GRCh37;
setting a path does not make GRCh38 intervals compatible.
For ORA, set both GMT_DIR (shell download location) and TRACECB_GMT_DIR
(Python reader location) to the same directory. Set RESULT_ROOT for the ORA
wrapper and provide its study/GTF inputs. Neither this wrapper nor the S-LDSC
wrapper automatically sources scripts/config.sh.
7. Record what another researcher needs
Alongside each run, save the code commit and local diff, resolved configuration without credentials, software/environment versions, data source/access details, genome builds, input checksums, simulation seeds and exact commands. Include a manifest connecting each paper figure to its input tables and producing command. Keep API tokens out of config files committed to Git. Complete path configuration is necessary, but does not replace missing inputs, scientific validation or the provenance required for exact paper reproduction.