A computational framework for local inference, fine-tuning, and evaluation of Large Language Models (LLMs) specialized for biological and therapeutic applications.
Systematically benchmark a LLM's ability to interpret genetic variants (ACMG/AMP guidelines) from genes and mutation and expression molecular profiles. Specifically to predict if over/under expression of a gene and a co-occurring mutation would be [Benign, Likely benign, Likely pathogenic, Pathogenic, Uncertain significance].
LLM_therapeutic provides a streamlined pipeline for running open-weight language models (e.g, Qwen2.5, TxGemma) on domain-specific biomedical tasks. This repository supports quantized local inference, token-level evaluation, and custom dataset preparation for bioinformatics workflows.
- Quantized Local Inference: Standardized
llama.cpp/ GGUF execution pipelines optimized for local setups with limited resources. - Biomedical Tokenization & Data Pipelines: Custom preprocessing modules for multi-omics metadata, gene expression profiles and mutational variants.
- Evaluation Benchmarks: Scripts to evaluate model outputs against published clinical data.
- Parameter-efficient fine-tuning: Tune foundation model with molecular data for cancer variant prediction and evaluate performance improvement.
- Applying parameter-efficient fine-tuning (PEFT) to human gene profiles (sub-sample of genes) improved model sensitivity for high-risk variants, raising both balanced accuracy and recall by 0.11 for samples with over-expressed genes with co-occurring mutations. While this resulted in a marginal precision loss (0.02), the net performance gain was substantial. Final clinical utility will depend on whether downstream applications prioritize minimizing false negatives over maintaining higher precision.
Clone the repository and set up a virtual environment
git clone [https://github.com/jordan2lee/LLM_therapeutic.git](https://github.com/jordan2lee/LLM_therapeutic.git)
cd LLM_therapeutic
python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txtMake sure to carry over the submodule
git submodule init
git submodule updateInstall llama.cpp by building from source
cd src
git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build --config ReleaseThen add this to PATH. Example export PATH="$HOME/LLM_therapeutic/src/llama.cpp/build/bin:$PATH" in .zshrc
Next locally download the model, will initially connect to the internet to download, cache and save model on local drive. After that will run entirely local. Using the Qwen model family because it performs well on medical, scientific, and biological benchmarks. No hf login needed for public GGUF mirror
hf download paultimothymooney/Qwen2.5-7B-Instruct-Q4_K_M-GGUF \
qwen2.5-7b-instruct-q4_k_m.gguf \
--local-dir modelsThis set provides a fully local LLM with inference on-device so nothing is sent to external servers (proprietary code, patient identifiers, creds, etc)
Using an external public dataset, determine which genes are mutated and gene expression profiles (differential gene expression). Feed these genes and several controls (not mutated, normal expression, not mutated high expression, etc) into LLM for predictions on biological relevance. Then use the external dataset to benchmark predictions from LLM.
Use the public cleaned TCGA data that is described in the Nature paper (Ellrott et al 2025) and is referenced in the Tumor Molecular Pathology Toolkit GitHub
wget -P src https://api.gdc.cancer.gov/data/5116e86f-7646-4b7b-9d6e-dafddf2cc0f3Then decompress file tar -xf TMP_20230209.tar.gz
python scripts/build_ref.pyTo change default files use -i and -o
Programmatically query ClinVar (public literature and other sources) for clinical significance of these genes based on mutation status and gene expression profile
submodule/clinvar-genes/scripts/clinvar_genes.py submodule/clinvar-genes/tests/fixtures/gene_list.txt submodule/clinvar-genes/results.ndjson > results/clinical_true.tsvFile clinical_true.tsv will be used to assess LLM performance downstream
By default it does not have permission to write to disk. You can enable this or simply have the outputs send to standard out and save those manually as a file. Shown is the second way.
bash scripts/get_prompts.shThen save the stdout to a file called results/prompts.txt
This model is run locally, so it is hardcoded to up the tokens. DO NOT modify to run on the cloud unless you decrease the tokens are check the projected cost to run.
bash scripts/llm_testing.shThen save the stdout to a file called results/responses_LLM.txt
File responses_LLM.txt will be used to benchmark against public peer-reviewed literature and other data sources
Consolidate results from different files into a single summary table. Then assess how well the model captures true clinical attributes.
python scripts/build_summary.py --outfile results/summary.tsv
python scripts/benchmark.py --inputfile results/summary.tsvUse LoRA to perform the PEFT to freeze original model weights and train low-rank adapters that will be put into the attention and feed-forward (MLP) layers.
GGUF is a main inference format for llama.cpp and so will start with the original Qwen model (Qwen2.5-7B-Instruct in 16-bit/Safetensors form)
hf auth login then follow prompts
Download model
hf download Qwen/Qwen2.5-7B-Instruct \
--local-dir models/qwen2.5-7b-instructRun parameter-efficient fine-tuning of LLM by applying LoRA (low-rank adaptation). This means the base LLM weights are frozen and only the newly injected adapter is trained.
python scripts/fine_tune_lora.py --train_file results/train.jsonl \
--base_model models/qwen2.5-7b-instruct \
--output_dir models/qwen2.5-7b-variant-loraRun newly fine-tuned model
bash scripts/infer.shThen save the stdout to a file called results/responses_PEFT_LLM.txt
File responses_PEFT_LLM.txt will be used to benchmark against base LLM
Consolidate results from different files into a single summary table. Then assess how well the model captures true clinical attributes.
python scripts/build_summary.py \
--resp_file results/responses_PEFT_LLM.txt \
--outfile results/summary_PEFT.tsv
python scripts/benchmark.py --inputfile results/summary_PEFT.tsvEllrott et al. (2025). Classification of non-TCGA cancer samples to TCGA molecular subtypes using compact feature sets. Cancer Cell, 43(2), 195–212.e11. https://doi.org/10.1016/j.ccell.2024.12.002
Tumor Molecular Pathology Toolkit: NCICCGPO/gdan-tmp-models
Landrum MJ, Lee JM, Riley GR, Jang W, Rubinstein WS, Church DM, Maglott DR. ClinVar: public archive of relationships among sequence variation and human phenotype. Nucleic Acids Res. 2014 Jan;42(Database issue):D980-5. doi: 10.1093/nar/gkt1113. Epub 2013 Nov 14. PMID: 24234437; PMCID: PMC3965032.