This document briefly describes how to contribute to DeepLC.
If you have an idea for a feature, use case to add or an approach for a bugfix, it is best to communicate with the community by creating an issue in GitHub issues.
DeepLC predicts peptide retention times using a 1D-CNN model trained on atomic composition features.
- Feature extraction (
deeplc/_features.py) —encode_peptidoform()converts a ProForma peptidoform string into per-position atomic composition arrays (C, H, N, O, S, P) plus a 20-AA one-hot encoding. Padding to 60 residues. Usespsm_utils.Peptidoformas the canonical peptide representation. - Dataset (
deeplc/data.py) —DeepLCDatasetwraps lists ofPeptidoformobjects into a PyTorchDataset. Features are encoded lazily in__getitem__.split_datasets()handles train/validation splits. - Model architecture (
deeplc/_architecture.py) —DeepLCModel: Conv1D branches (atomic, summed-atomic, global, one-hot) feed a shared dense trunk intoBatchedHeadsreturning[batch, n_heads]. An optional fine-tuning adapter (self.adapter, attached viaadd_adapter()) maps head output to[batch, 1]. - Training/inference (
deeplc/_model_ops.py) —load_model(),train(),predict(),evaluate(). Checkpoints are plain state dicts loaded withweights_only=True. - Calibration (
deeplc/calibration/package) —predict(..., return_matrix=True)always returns a(n, n_heads)matrix,n_heads=1for a single-setup model.deeplc/calibration/simple.pyholds naive, single-series calibrations (CalibrationABC,IdentityCalibration,PiecewiseLinearCalibration,SplineTransformerCalibration) that map one column onto observed RT space and know nothing about heads.deeplc/calibration/multihead.pyholdsMultiHeadCalibrationABC and the classes that pick their own head(s) from the matrix:MultiHeadPiecewiseLinearCalibrationandMultiHeadSplineCalibration(rank heads by correlation, delegate to the matchingsimple.pyclass) andMultiHeadRidgeCalibration(combines the best-correlating heads with a ridge fit).core.calibrate()/predict_and_calibrate()only acceptMultiHeadCalibrationinstances and always hand over the full matrix; the default isMultiHeadRidgeCalibration. - Reference selection (
deeplc/_reference_selection.py) — selects high-confidence PSMs from input for auto-calibration. - Core API (
deeplc/core.py) — top-level functions:predict(),calibrate(),predict_and_calibrate(),finetune_and_predict(). - CLI (
deeplc/__main__.py) — two subcommands,predictandgui. Reads PSM files viapsm_utils.io.read_file(). - GUI (
deeplc/gui.py) — NiceGUI-based web UI, launchable as browser app or native desktop window viapywebview.
Conventions: peptide sequences use ProForma notation throughout (via psm_utils.Peptidoform);
PSM collections are psm_utils.PSMList; line length is 99 characters (ruff); Python >= 3.11 with
from __future__ import annotations.
- Fork DeepLC on GitHub to make your changes.
- Commit and push your changes to your fork.
- Open a
pull request
with these changes. You pull request message ideally should include:
- A description of why the changes should be made.
- A description of the implementation of the changes.
- A description of how to test the changes.
- The pull request should pass all the continuous integration tests which are automatically run by GitHub Actions.
-
When a new version is ready to be published:
- Change the version number in
setup.pyusing semantic versioning. - Update the changelog (if not already done) in
CHANGELOG.mdaccording to Keep a Changelog. - Set a new tag with the version number, e.g.
git tag v0.1.5. - Push to GitHub, with the tag:
git push; git push --tags.
- Change the version number in
-
When a new tag is pushed to (or made on) GitHub that matches
v*, the following GitHub Actions are triggered:- The Python package is build and published to PyPI.
- A zip archive is made of the
./deeplc_gui/directory, excluding./deeplc_gui/srcwith Zip Release. - A GitHub release is made with the zipped GUI files as assets and the new
changes listed in
CHANGELOG.mdwith Git Release. - After some time, the bioconda package should get updated automatically.