Skip to content
shadi97khPublic

About

Biology-informed regularization and a perturbation-based saliency faithfulness protocol for siRNA efficacy prediction.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

300 Commits

Folders and files

Repository files navigation

BioPrior

python >3.8 pytorch

Validating Interpretability in siRNA Efficacy Prediction: A Perturbation-Based, Dataset-Aware Protocol

BioPrior extends OligoFormer with biology-informed regularization and a perturbation-based saliency validation protocol for siRNA efficacy prediction. We introduce a pre-synthesis gate that tests whether gradient-based saliency maps faithfully identify model-sensitive positions before they are used to guide sequence design decisions.

📄 Paper: arXiv 👥 Authors: Zahra Khodagholi, Niloofar Yousefi — University of Central Florida

Pipeline Overview Positioning saliency validation in the lab-in-the-loop decision pipeline. Our protocol acts as a pre-synthesis gate before saliency maps are used for design guidance.

Key Contributions

  1. Saliency validation protocol — A composition-controlled, perturbation-based faithfulness test with nucleotide-matched baselines and pass/fail criteria for deploy-time decisions.
  2. Cross-dataset faithfulness analysis — 19/20 fold–dataset settings pass; cross-dataset transfer reveals two failure modes (faithful-but-wrong and inverted saliency) that would otherwise go undetected.
  3. Biology-informed regularization (BioPrior) — Differentiable constraints encoding thermodynamic asymmetry, seed region composition, GC heuristics, and immune motif avoidance.
  4. Protocol-level diagnostic — Taka (luciferase reporter) exhibits systematic incompatibility with mRNA-level assay datasets, demonstrating that assay shifts can silently invalidate explanations.

Method Overview Counterfactual Faithfulness Top: Training with BioPrior constraints. Bottom: Saliency validation via expected-effect perturbations with nucleotide-matched baselines.

Datasets

We evaluate on four benchmark groups spanning distinct experimental protocols and cell lines:

Dataset siRNA Number Cell Line Group
Huesken 2,431 H1299 Hu
Reynolds 240 HEK293 Mix
Vickers 76 T24 Mix
Harborth 44 HeLa Mix
Ui-Tei 62 HeLa Mix
Khvorova 14 HEK293 Mix
Hsieh 108 HEK293T Mix
Amarzguioui 46 Cos-1, HaCaT Mix
Takayuki 702 HeLa Taka
Shabalina 653 Multiple Shabalina

Following OligoFormer, we organize the ten studies above into four evaluation groups: Hu (Huesken alone), Mix (the seven studies labeled Mix in the table — Reynolds, Vickers, Harborth, Ui-Tei, Khvorova, Hsieh, Amarzguioui — totaling 581 siRNAs after deduplication), Taka (Takayuki), and Shabalina. Cross-dataset transfer experiments reveal that Taka (luciferase reporter assay, single target) is systematically incompatible with the other three groups.

Installation

Requirements

  • Python 3.8+
  • PyTorch 1.12+
  • CUDA 11.3+ (NVIDIA GPU with ~8GB memory)
  • ViennaRNA (for thermodynamic features)

Setup

# Clone repository
git clone https://github.com/shadi97kh/BioPrior.git
cd BioPrior

# Create environment
conda create -n bioprior python=3.8
conda activate bioprior

# Install PyTorch (adjust CUDA version as needed)
conda install pytorch==1.12.0 torchvision torchaudio cudatoolkit=11.3 -c pytorch

# Install dependencies
pip install -r requirements.txt

# Install ViennaRNA for thermodynamic calculations
conda install -c bioconda viennarna

RNA-FM Embeddings

RNA-FM embeddings are required for sequence representation:

# Option 1: Download pre-packaged RNA-FM
wget https://cloud.tsinghua.edu.cn/f/46d71884ee8848b3a958/?dl=1 -O RNA-FM.tar.gz
tar -zxvf RNA-FM.tar.gz

# Option 2: Clone and set up RNA-FM from source
git clone https://github.com/ml4bio/RNA-FM.git
cd RNA-FM
conda env create --name RNA-FM -f environment.yml
conda activate RNA-FM

Download pre-trained RNA-FM weights from this gdrive link and place .pth files into the pretrained folder.

# Generate RNA-FM features for all datasets
conda activate RNA-FM
bash scripts/RNA-FM-features.sh
conda activate bioprior

Verify Installation

# Quick test — should print model summary and exit
python scripts/main.py --datasets Hu --epoch 1 --cuda 0

Usage

Intra-Dataset Evaluation (5-fold CV)

Train and evaluate with BioPrior on individual datasets:

# Huesken dataset
python scripts/main.py --datasets Hu --val_mode intra --epoch 300 --early_stopping 50 --cuda 0

# Takayuki dataset
python scripts/main.py --datasets Taka --val_mode intra --epoch 300 --early_stopping 50 --cuda 0

# Mix dataset
python scripts/main.py --datasets Mix --val_mode intra --epoch 300 --early_stopping 50 --cuda 0

# Shabalina dataset
python scripts/main.py --datasets Shabalina --val_mode intra --epoch 300 --early_stopping 50 --cuda 0

Cross-Dataset Transfer

Train on one dataset, evaluate on another:

# Train on Hu, test on Mix
python scripts/main.py --datasets Hu Mix --val_mode inter --epoch 300 --early_stopping 50 --cuda 0

# Train on Hu, test on Taka
python scripts/main.py --datasets Hu Taka --val_mode inter --epoch 300 --early_stopping 50 --cuda 0

# Train on Mix, test on Hu
python scripts/main.py --datasets Mix Hu --val_mode inter --epoch 300 --early_stopping 50 --cuda 0

# Train on Shabalina, test on Mix
python scripts/main.py --datasets Shabalina Mix --val_mode inter --epoch 300 --early_stopping 50 --cuda 0

# Train on Taka, test on Hu (demonstrates inverted saliency)
python scripts/main.py --datasets Taka Hu --val_mode inter --epoch 300 --early_stopping 50 --cuda 0

Ablation: Baseline without BioPrior

The --ablation flag controls the regularization mode: baseline (no physics), mechanistic (default, full BioPrior), full, or neutral.

# Disable biology-informed regularization
python scripts/main.py --datasets Hu --val_mode intra --epoch 300 --early_stopping 50 --cuda 0 --ablation baseline

Saliency Faithfulness Validation

Note: The model/ directory is not committed. Train a model first (commands above) — checkpoints will be written to ./model/best_model.pth by default — then point --model_path at the resulting file.

Run the perturbation-based faithfulness test on a trained model. Each run reports top-k effect, bottom-k effect, and the nucleotide-matched random baseline — the bottom-k and matched-baseline contrasts are computed automatically as built-in negative controls.

# Intra-dataset faithfulness (Hu, default k=3, 50 matched samples)
python scripts/perturbation_test.py --dataset Hu --model_path model/best_model.pth --k_positions 3 --n_matched 50 --cuda 0

# Vary k for sensitivity analysis
python scripts/perturbation_test.py --dataset Hu --model_path model/best_model.pth --k_positions 1 --cuda 0
python scripts/perturbation_test.py --dataset Hu --model_path model/best_model.pth --k_positions 5 --cuda 0

# Cross-dataset transfer faithfulness (train on Hu, test saliency on Taka)
python scripts/perturbation_test.py --dataset Taka --model_path model/hu_best_model.pth --k_positions 3 --cuda 0

# Inverted saliency demo (train on Taka, test saliency on Hu)
python scripts/perturbation_test.py --dataset Hu --model_path model/taka_best_model.pth --k_positions 3 --cuda 0

Saliency Visualization

# Generate saliency maps and position-importance plots
python scripts/saliency_analysis/saliency_maps.py

Additional analyses live in scripts/saliency_analysis/ (5'-end swap experiment, similarity analyses, etc.).

Model Inference on New Sequences

# Input mRNA fasta (traverse with 19nt window)
python scripts/main.py --infer 1 -i1 data/example.fa --cuda 0

# Input mRNA + specific siRNAs
python scripts/main.py --infer 1 -i1 data/example.fa -i2 data/example_siRNA.fa --cuda 0

# Manual mRNA input
python scripts/main.py --infer 2 --cuda 0

Results Summary

Intra-Dataset Performance (5-fold CV, +BioPrior)

Dataset AUC PR-AUC PCC F1
Hu 0.83 ± .02 0.82 ± .02 0.64 ± .02 0.77 ± .03
Mix 0.81 ± .05 0.81 ± .07 0.61 ± .08 0.77 ± .06
Taka 0.84 ± .05 0.67 ± .10 0.68 ± .08 0.60 ± .07
Shabalina 0.72 ± .02 0.66 ± .06 0.49 ± .03 0.68 ± .03

Saliency Faithfulness (19/20 fold–dataset combinations pass)

Dataset Win % Cohen's d_z Status
Hu 85.2 ± 4.1 0.86 ± .26 ✓
Mix 83.7 ± 6.2 0.93 ± .45 ✓
Taka 87.1 ± 3.8 1.07 ± .24 ✓
Shabalina 81.4 ± 12.3 0.70 ± .42 ✓†

†4/5 folds pass; 1 fold shows inverted saliency (d_z = −1.20).

Cross-Dataset Transfer: Two Failure Modes

Taka Anomaly The Taka transfer anomaly: (a) Taka-trained models degrade during training when tested on other datasets. (b) All source models fail on Taka. (c) Successful transfers converge at intermediate epochs while Taka pairs never learn.

Failure Mode Example AUC Win % d_z
Faithful-but-wrong Mix → Taka 0.497 97.5 1.47
Inverted saliency Taka → Hu 0.490 9.5 −1.25

Project Structure

BioPrior/
├── data/
│   ├── Hu.csv                          # Huesken dataset (2,431 siRNAs)
│   ├── Mix.csv                         # Mix dataset (581 siRNAs, 7 studies)
│   ├── Taka.csv                        # Katoh/Takayuki dataset (702 siRNAs)
│   ├── Shabalina.csv                   # Shabalina dataset (653 siRNAs)
│   ├── example.fa                      # Example mRNA fasta
│   └── example_siRNA.fa                # Example siRNA fasta
├── figures/                            # Paper figures
├── scripts/
│   ├── main.py                         # Training and evaluation entry point
│   ├── train.py / train_single.py      # Training loops (k-fold and single-split)
│   ├── test.py  / test_single.py       # Evaluation entry points
│   ├── model.py                        # Model architecture (Conv-BiLSTM-Transformer)
│   ├── learnable_physics_loss.py       # BioPrior regularization module
│   ├── perturbation_test.py            # Saliency faithfulness validation
│   ├── saliency_analysis/              # Saliency maps + downstream analyses
│   ├── loader.py / loader_infer.py     # Dataset loading and preprocessing
│   ├── infer.py                        # Inference on new sequences
│   ├── metrics.py                      # Evaluation metrics
│   ├── logger.py / train_logger.py     # Training loggers
│   ├── flanking_mRNA_asymmetric.py     # Asymmetric mRNA context
│   ├── flanking_mRNA_symmetric.py      # Symmetric mRNA context
│   ├── mismatch.py                     # Mismatch handling
│   ├── TD_calculation.ipynb            # Thermodynamic-feature notebook
│   └── RNA-FM-features.sh              # RNA-FM feature generation
├── environment.yml
├── requirements.txt
└── README.md

Trained checkpoints are written to ./model/ and off-target reference files (if you enable --off_target) are expected under ./off-target/ref/. Neither directory is committed.

Acknowledgments

This project builds on OligoFormer by Bai et al. (2024). We thank the original authors for making their code publicly available.

Citations

If you use BioPrior in your research, please cite:

@inproceedings{
khodagholi2026validating,
title={{VALIDATING} {INTERPRETABILITY} {IN} {SIRNA} {EFFICACY} {PREDICTION}: A {PERTURBATION}-{BASED}, {DATASET}- {AWARE} {PROTOCOL}},
author={Shadi Khodagholi and Niloofar Yousefi},
booktitle={ICLR 2026 Workshop on Machine Learning for Genomics Explorations},
year={2026},
url={https://openreview.net/forum?id=0WStQt1qw5}
}

And the original OligoFormer paper:

@article{bai2024oligoformer,
  title={OligoFormer: an accurate and robust prediction method for siRNA design},
  author={Bai, Yilan and Zhong, Haochen and Wang, Taiwei and Lu, Zhi John},
  journal={bioRxiv},
  pages={2024--02},
  year={2024},
  publisher={Cold Spring Harbor Laboratory}
}

License

This project is for academic and non-commercial use. The original OligoFormer code is subject to its own license terms.

About

Biology-informed regularization and a perturbation-based saliency faithfulness protocol for siRNA efficacy prediction.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages