Skip to content

About

Implementing and tuning Large Language Models for variant applications

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

38 Commits

Folders and files

Repository files navigation

LLM Therapeutic: Variant Interpretation & Benchmarking Pipeline

A computational framework for local inference, fine-tuning, and evaluation of Large Language Models (LLMs) specialized for biological and therapeutic applications.

Systematically benchmark a LLM's ability to interpret genetic variants (ACMG/AMP guidelines) from genes and mutation and expression molecular profiles. Specifically to predict if over/under expression of a gene and a co-occurring mutation would be [Benign, Likely benign, Likely pathogenic, Pathogenic, Uncertain significance].

Overview

LLM_therapeutic provides a streamlined pipeline for running open-weight language models (e.g, Qwen2.5, TxGemma) on domain-specific biomedical tasks. This repository supports quantized local inference, token-level evaluation, and custom dataset preparation for bioinformatics workflows.

Key Features

  • Quantized Local Inference: Standardized llama.cpp / GGUF execution pipelines optimized for local setups with limited resources.
  • Biomedical Tokenization & Data Pipelines: Custom preprocessing modules for multi-omics metadata, gene expression profiles and mutational variants.
  • Evaluation Benchmarks: Scripts to evaluate model outputs against published clinical data.
  • Parameter-efficient fine-tuning: Tune foundation model with molecular data for cancer variant prediction and evaluate performance improvement.

Results Preview

  • Applying parameter-efficient fine-tuning (PEFT) to human gene profiles (sub-sample of genes) improved model sensitivity for high-risk variants, raising both balanced accuracy and recall by 0.11 for samples with over-expressed genes with co-occurring mutations. While this resulted in a marginal precision loss (0.02), the net performance gain was substantial. Final clinical utility will depend on whether downstream applications prioritize minimizing false negatives over maintaining higher precision.

Quickstart

1. Installation

Clone the repository and set up a virtual environment

git clone [https://github.com/jordan2lee/LLM_therapeutic.git](https://github.com/jordan2lee/LLM_therapeutic.git)
cd LLM_therapeutic

python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt

Make sure to carry over the submodule

git submodule init
git submodule update

2. Environment Setup

Install llama.cpp by building from source

cd src
git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build --config Release

Then add this to PATH. Example export PATH="$HOME/LLM_therapeutic/src/llama.cpp/build/bin:$PATH" in .zshrc

Next locally download the model, will initially connect to the internet to download, cache and save model on local drive. After that will run entirely local. Using the Qwen model family because it performs well on medical, scientific, and biological benchmarks. No hf login needed for public GGUF mirror

hf download paultimothymooney/Qwen2.5-7B-Instruct-Q4_K_M-GGUF \
    qwen2.5-7b-instruct-q4_k_m.gguf \
    --local-dir models

This set provides a fully local LLM with inference on-device so nothing is sent to external servers (proprietary code, patient identifiers, creds, etc)

1. Construct and Benchmark LLM

Using an external public dataset, determine which genes are mutated and gene expression profiles (differential gene expression). Feed these genes and several controls (not mutated, normal expression, not mutated high expression, etc) into LLM for predictions on biological relevance. Then use the external dataset to benchmark predictions from LLM.

Build Reference Dataset

Use the public cleaned TCGA data that is described in the Nature paper (Ellrott et al 2025) and is referenced in the Tumor Molecular Pathology Toolkit GitHub

wget -P src https://api.gdc.cancer.gov/data/5116e86f-7646-4b7b-9d6e-dafddf2cc0f3

Then decompress file tar -xf TMP_20230209.tar.gz

python scripts/build_ref.py

To change default files use -i and -o

Programmatically query ClinVar (public literature and other sources) for clinical significance of these genes based on mutation status and gene expression profile

submodule/clinvar-genes/scripts/clinvar_genes.py submodule/clinvar-genes/tests/fixtures/gene_list.txt submodule/clinvar-genes/results.ndjson > results/clinical_true.tsv

File clinical_true.tsv will be used to assess LLM performance downstream

Run Base LLM

Get prompts for LLM testing

By default it does not have permission to write to disk. You can enable this or simply have the outputs send to standard out and save those manually as a file. Shown is the second way.

bash scripts/get_prompts.sh

Then save the stdout to a file called results/prompts.txt

Get LLM calls

This model is run locally, so it is hardcoded to up the tokens. DO NOT modify to run on the cloud unless you decrease the tokens are check the projected cost to run.

bash scripts/llm_testing.sh

Then save the stdout to a file called results/responses_LLM.txt

File responses_LLM.txt will be used to benchmark against public peer-reviewed literature and other data sources

Benchmark Performance

Consolidate results from different files into a single summary table. Then assess how well the model captures true clinical attributes.

python scripts/build_summary.py --outfile results/summary.tsv
python scripts/benchmark.py --inputfile results/summary.tsv

2. Parameter-Efficient Fine-Tuning

Use LoRA to perform the PEFT to freeze original model weights and train low-rank adapters that will be put into the attention and feed-forward (MLP) layers.

GGUF is a main inference format for llama.cpp and so will start with the original Qwen model (Qwen2.5-7B-Instruct in 16-bit/Safetensors form)

hf auth login then follow prompts

Download model

hf download Qwen/Qwen2.5-7B-Instruct \
    --local-dir models/qwen2.5-7b-instruct

Run parameter-efficient fine-tuning of LLM by applying LoRA (low-rank adaptation). This means the base LLM weights are frozen and only the newly injected adapter is trained.

python scripts/fine_tune_lora.py --train_file results/train.jsonl \
    --base_model models/qwen2.5-7b-instruct \
    --output_dir models/qwen2.5-7b-variant-lora

Run newly fine-tuned model

bash scripts/infer.sh

Then save the stdout to a file called results/responses_PEFT_LLM.txt

File responses_PEFT_LLM.txt will be used to benchmark against base LLM

Benchmark Performance

Consolidate results from different files into a single summary table. Then assess how well the model captures true clinical attributes.

python scripts/build_summary.py \
    --resp_file results/responses_PEFT_LLM.txt \
    --outfile results/summary_PEFT.tsv
python scripts/benchmark.py --inputfile results/summary_PEFT.tsv

Citations & References

Ellrott et al. (2025). Classification of non-TCGA cancer samples to TCGA molecular subtypes using compact feature sets. Cancer Cell, 43(2), 195–212.e11. https://doi.org/10.1016/j.ccell.2024.12.002

Tumor Molecular Pathology Toolkit: NCICCGPO/gdan-tmp-models

Landrum MJ, Lee JM, Riley GR, Jang W, Rubinstein WS, Church DM, Maglott DR. ClinVar: public archive of relationships among sequence variation and human phenotype. Nucleic Acids Res. 2014 Jan;42(Database issue):D980-5. doi: 10.1093/nar/gkt1113. Epub 2013 Nov 14. PMID: 24234437; PMCID: PMC3965032.

About

Implementing and tuning Large Language Models for variant applications

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages