User guide Database navigation Structure, function and target discovery

HANSEN: A detailed guide to using HANSEN

This guide explains how to search HANSEN, interpret protein pages, use the structure viewers, explore GO annotations, inspect model statistics and move from database-wide summaries to detailed protein-level evidence.

Fast Routes
Best starting point: use the home-page query panel if you know a UniProt accession, ML locus tag, gene, protein name, ligand code or sequence fragment.

1. What HANSEN contains

HANSEN is an integrated structural and functional resource for the Mycobacterium leprae proteome. It brings together core protein identifiers, curated functional annotations, sequence features, cross-references, homology and de novo structure models, oligomeric assemblies, ligand annotations, pocket predictions, AF2Bind binding-site predictions, PAE confidence maps and B-cell epitope propensity information.

UniProt and ML locus identifiers Gene and protein annotation AF3 / Boltz / Chai / Boltz-2 models Monomers and oligomers Mol* structure viewer pLDDT and PAE Pockets and AF2Bind Ligands and PubChem links GO browser Database statistics

2. Home page and search

The home page is the main entry point. Use the query panel to search by identifier, name, ligand or sequence. The top navigation bar provides direct access to the query area, database background, statistics dashboard and GO browser.

Use the query panel when you know a target

Search using a UniProt accession, ML locus tag, gene name, protein name, ligand symbol, ligand name, PubChem CID or protein sequence fragment.

Use global pages when you are exploring

Open the Statistics dashboard for database-wide coverage or the GO Browser to discover groups of proteins by ontology.

About the resource

About HANSEN and database scope

This section summarises the resource scope, while the home page remains focused on search and navigation.

About HANSEN

HANSEN supports translational and mechanistic leprosy research by making M. leprae protein-centric information easy to search, inspect and reuse. The platform links identifiers, functional evidence, sequence features and structural models in a format suitable for target triage, biomarker assessment, comparative interpretation and downstream experimental design.

Targeted retrieval
Search M. leprae proteins by UniProt entry, ML locus tag, gene name, protein name, ligand, or full or partial amino-acid sequence.
Scientifically grounded summaries
Protein-centric pages provide concise, biologically meaningful summaries of annotation, structure, localisation, sequence features and linked evidence relevant to leprosy research.
Leprosy-relevant annotation
GO terms, pathway context, sequence features, domain architecture and supporting metadata are presented for direct interpretation of M. leprae biology and disease relevance.
Structure-enabled interpretation
Monomeric and oligomeric models, curated sequences and browsable tables are organised for downstream analysis in target assessment, comparative genomics and diagnostic assay development.

HANSEN scope


Use HANSEN to interrogate M. leprae proteins linked to metabolism, cell-envelope biology, host interaction, persistence, biomarker discovery and structure-guided target prioritisation across the leprosy bacillus proteome.

4. Query types

Query type What to enter What HANSEN does Best use
UniProt UniProt accession or entry identifier Opens the matching protein page. When you know the canonical UniProt entry.
ML locus tag Example: ML0005 Finds the protein by M. leprae locus tag. Best for genome and proteome workflows.
Gene name Gene symbol or partial gene text Uses exact, prefix and partial-text matching. When working from annotation tables or literature.
Protein name Full or partial protein description Finds the closest matching annotated protein name. When you know the biological function but not the locus tag.
Sequence Full or partial amino-acid sequence Normalises the sequence and searches for exact, contained and local matches. Useful for fragments, peptides or copied FASTA sequence.
Ligand Ligand code, name, SMILES-derived annotation, or PubChem CID Returns proteins linked to that ligand through imported oligomer/template ligand annotations. Useful for identifying ligand-associated targets or cofactor-binding proteins.

5. Protein result page

A protein result page is the main detailed view for a single HANSEN entry. It usually begins with a summary panel containing the UniProt entry, ML locus tag, gene name, protein name, organism, length and other core identifiers. The rest of the page is divided into functional, structural and evidence sections.

ASummary and identifiers

Use this section to confirm that you have opened the correct protein. It contains entry names, gene names, ML locus tags and links to external resources where available.

BFunctional annotation

Review EC numbers, catalytic activity, cofactors, pathways, binding sites, active sites, protein family, domains, motifs and other curated annotations.

CEvidence and cross-references

Check PubMed IDs, DOI IDs, InterPro, KEGG, GeneID, STRING and other cross-references to connect HANSEN annotations with external evidence.

DStructure and target-discovery panels

Inspect model quality, Mol* structures, PAE maps, pockets, AF2Bind predictions, ligands and epitope predictions to prioritise proteins.

6. Structure models

The model section lists the structural models available for the selected protein. Depending on data availability, HANSEN can show monomeric and oligomeric models generated by AF3, Boltz, Chai and Boltz-2.

Model type Meaning How to use it
Monomer Single-protein structural model. Best for inspecting domain architecture, local residues, active sites and local confidence.
Homomer Oligomer made from repeated copies of the same ML protein. Use to inspect biological assemblies, repeated interfaces and symmetry-related pockets.
Heteromer Complex containing different protein components. Use to inspect protein-protein interfaces, complex-specific pockets and chain-level behaviour.
Template-linked oligomer Assembly reconstructed or modelled using template evidence. Use template PDB and assembly metadata to interpret biological relevance.
Interpretation note: a model is a prediction unless experimental structural evidence is explicitly indicated. Use pLDDT, PAE, pocket evidence and biological annotation together before prioritising a target.

7. Mol* viewer

The Mol* panel is the interactive 3D structure viewer. Use it to rotate and zoom the model, inspect chains, focus on residues or ligands, and view pocket or AF2Bind overlays when available.

  • Load or select a model: choose a structure model from the model controls. Oligomers may take longer to load than monomers.
  • Rotate and zoom: drag to rotate, scroll to zoom and right-click or secondary-drag to pan, depending on your device.
  • Use full-page mode: open the larger Mol* view when detailed inspection is needed.
  • Focus features: use ligand, pocket, AF2Bind or epitope controls to focus the relevant region in 3D.
  • Return to the page: exit full-page mode to continue reviewing tables and annotation cards.

8. pLDDT and PAE confidence interpretation

Confidence panels help distinguish reliable local structural regions from lower-confidence or flexible areas. pLDDT reports residue-level confidence, while PAE helps assess the reliability of domain-domain or chain-chain placement.

pLDDT

Use pLDDT to judge local residue confidence. High-confidence residues are more reliable for interpreting local pockets, active sites and epitopes. Low-confidence regions may be flexible, disordered or modelled with greater uncertainty.

PAE

Use PAE to judge relative placement of domains and chains. For oligomers, inspect block patterns across chains to understand interface confidence and assembly reliability.

For large oligomers, load the structure first and load the PAE map only when needed. PAE files can be large and may slow the initial page rendering.

9. Pockets, hotspots and AF2Bind

HANSEN includes predicted pockets and AF2Bind-style binding-site predictions where available. These panels help prioritise proteins and residues for druggability assessment.

  • Pocket table: lists predicted pockets, scores, tools and residue information where available.
  • Focus pocket: centres Mol* on the selected pocket and overlays the pocket region if residue-level data are available.
  • AF2Bind: highlights predicted binding residues and can be used alongside pocket predictions.
  • Remove/hide overlays: clear surfaces after inspection to return to a clean structure view.

For target prioritisation, favour proteins whose predicted pockets overlap with high-confidence structural regions, biologically relevant ligands, catalytic sites, conserved residues or essential functional annotations.

10. Ligands

Ligand annotations help connect modelled oligomers and template assemblies with cofactors, ions, substrates or other small molecules. Ligands may appear as CCD-style symbols, names, SMILES-derived entries or PubChem-linked compounds.

  • Search by code: examples include short ligand symbols such as ZN or template ligand codes.
  • Search by name: enter a ligand or compound name if the imported annotation includes it.
  • Search by PubChem CID: use formats such as CID 32051 when PubChem resolution is available.
  • Protein page ligand table: use ligand rows to focus the ligand in Mol* and open PubChem where linked.
If an expected reconstructed oligomer ligand is not shown, it may not yet be present in the ligand annotation table used by HANSEN. Ligand searches only include ligands that have been imported into the current annotation source.

11. B-cell epitope propensity

The B-cell epitope section shows residue-level or region-level epitope propensity where available. Use it to identify exposed, structurally plausible antigenic regions, especially when the predictions are considered alongside pLDDT, accessibility and 3D localisation.

  • Review the epitope table for predicted high-scoring residues or regions.
  • Focus predicted residues in Mol* to evaluate their structural context.
  • Compare epitope propensity with pockets, ligands and functional domains when prioritising diagnostic or immunological targets.

12. GO browser

The GO browser supports ontology-guided exploration. It uses a sunburst-style visualisation and associated protein tables to help users move from broad ontology categories to specific protein sets.

  • Hover: inspect a GO branch and see its context.
  • Mouse wheel: move inward or outward through GO hierarchy levels.
  • Click: lock the selected branch or term.
  • Review matched proteins: use the table below the sunburst to open individual protein pages.
  • Use namespace filters: compare biological process, molecular function and cellular component annotations.

GO browsing is useful when you do not know a specific protein target but want to explore all proteins annotated with a functional process, molecular activity or localisation.

13. Statistics dashboard

The Statistics page gives a database-wide summary of structural coverage, model confidence, assembly types, oligomeric states, ligand codes, pocket-score classes and epitope evidence.

Dashboard area What it tells you How to use it
Entry and annotation counts How many proteins have core annotation, GO terms, EC numbers or 3D structural information. Assess annotation completeness.
Model coverage How many proteins have monomeric, homomeric or heteromeric models. Identify modelling gaps and completed coverage.
pLDDT and PAE classes Confidence distribution across models. Prioritise high-confidence models for downstream interpretation.
Ligands Most frequent ligand annotations and linked proteins. Find cofactor-associated or ligand-template-associated proteins.
Pockets and epitopes Distribution of pocket predictions, pocket scores and epitope evidence. Shortlist targets for druggability or diagnostic antigen exploration.

14. Example workflows

  1. Go to Query, or try a direct example such as ML0005, gyrA, or ZN ligand search.
  2. Enter an ML locus tag, for example ML0005.
  3. Open the protein page and confirm the summary identifiers.
  4. Inspect model scores and select the most relevant monomer or oligomer.
  5. Use Mol* to inspect residues, ligands, pockets and chain interfaces.
  6. Load PAE if you need domain or chain-placement confidence.

  1. Use the Ligand query on the home page, or open an example such as ZN.
  2. Enter a ligand symbol, name or PubChem CID.
  3. Review the list of proteins linked to that ligand.
  4. Open each protein and inspect the ligand table and Mol* focus action.
  5. Combine ligand evidence with pockets, confidence and functional annotation.

  1. Open the GO Browser.
  2. Select a namespace or navigate through the sunburst.
  3. Click a GO branch or term to lock the protein set.
  4. Open candidate proteins from the table.
  5. Inspect structures, pockets, ligands and epitope predictions for each candidate.

  1. Start with Statistics to identify proteins with models, pockets and ligand annotations.
  2. Open proteins with high-confidence structures and biologically relevant assemblies.
  3. Prioritise proteins with strong pockets, AF2Bind residues, ligands or conserved functional sites.
  4. Check PAE for domain/interface reliability before interpreting oligomeric pockets.
  5. Use external links and annotations to assess biological relevance.

15. Troubleshooting and interpretation tips

Issue Likely reason What to do
Search gives no protein The query may not match the stored identifier or annotation text. Try UniProt, ML locus tag, gene name and partial protein name separately.
Sequence search gives no result The fragment may be too short, absent, or from a different strain/annotation set. Use a longer sequence fragment and remove spaces, numbers and FASTA headers.
Oligomer loads slowly Large CIF/BCIF and PAE files can be heavy. Load the structure first; load PAE only when needed.
PAE is missing The PAE JSON may not exist for that selected model or may not be mapped. Use another model or check whether the corresponding PAE file has been staged.
Pocket focus works but no surface appears Residue lists may be unavailable for that pocket source/model. Use available residue-level pocket outputs or rebuild pocket summaries with residues.
Ligand search misses expected ligands Reconstructed YAML ligands may not yet be imported into the ligand table. Ask the HANSEN maintainer to import the reconstructed ligand annotations, then refresh the page.
Recommended interpretation: do not prioritise a protein based on one signal alone. Combine functional annotation, biological relevance, model confidence, pocket evidence, ligand evidence, oligomeric context and experimental feasibility.

Target prioritisation ?

The Target Prioritisation page ranks Mycobacterium leprae proteins by integrating ProteomeLM-derived proteome-context information with structural, functional and druggability evidence.

Open Target Prioritisation

What the page shows

The table provides an evidence-weighted shortlist of candidate proteins for target discovery. Each row corresponds to an ML locus or protein and includes a Target Priority Score, priority tier, ProteomeLM contextual signal, pocket/AF2Bind evidence, model-quality evidence, annotation support and a short rationale explaining why the protein was ranked.

How the ProteomeLM signal was generated

Each protein is converted into an ESM-C (600M) embedding and passed through ProteomeLM-L in whole-proteome context, so it is interpreted relative to the rest of the M. leprae proteome rather than as an isolated sequence. A supervised ProteomeLM-Ess head, based on logistic regression of the contextual embeddings, is then trained on experimental M. tuberculosis transposon-sequencing (Tn-seq) essentiality labels (625 essential and 3,277 non-essential proteins), using mmseqs homology-grouped cross-validation. The trained model is transferred to score every M. leprae protein. In leak-free grouped cross-validation in M. tuberculosis, it reaches AUROC 0.84, compared with 0.81 for ESM-C and 0.71 for the earlier ESM-2 model. This ProteomeLM-Ess probability is now the primary ProteomeLM signal in the Target Priority Score.

Important interpretation: the ProteomeLM-Ess probability is an AI-derived, cross-species essentiality estimate transferred from M. tuberculosis experimental data — not a native, experimentally validated M. leprae essentiality measurement. High-ranking candidates still require manual biological review and experimental validation.

Evidence layers used

Evidence layer Role in prioritisation
ProteomeLM contextual signal AI-derived whole-proteome contextual evidence from ESM-C embeddings and ProteomeLM inference.
Boltz-2 model quality Supports confidence in structure-based interpretation, including pockets and residue-level signals.
Pocket evidence Captures predicted cavities and potential small-molecule binding sites.
AF2Bind evidence Summarises residue-level binding propensity and high-propensity binding regions.
Functional annotation Uses EC, GO, pathway, active-site, binding-site, cofactor and protein-family information.
Tractability features Considers properties such as protein size, enzymatic function and structural druggability support.

How the Target Priority Score is calculated

The final score is a composite 0–100 score combining AI-derived context, druggability, annotation, model quality and tractability:

Target Priority Score = 100 × ( 0.35 × ProteomeLM contextual score + 0.25 × pocket/AF2Bind binding-site score + 0.20 × annotation score + 0.10 × Boltz-2 model-quality score + 0.10 × tractability score )

Priority tiers

How to use the ranking

Use the table as a discovery shortlist rather than as a final decision. The most compelling targets are those with a high Target Priority Score, strong ProteomeLM contextual signal, confident Boltz-2 model, clear pocket or AF2Bind evidence, relevant pathway or enzyme annotation and biological plausibility.

Current limitation

HANSEN now provides a supervised ProteomeLM-Ess essentiality probability, trained on M. tuberculosis transposon-sequencing (Tn-seq) essentiality labels and transferred to M. leprae through the shared ProteomeLM contextual feature space. The remaining limitation is biological: these are cross-species predictions, not native experimental M. leprae knockout data. They should therefore guide, rather than replace, experimental target validation.

HANSEN essentiality prediction

This section explains how HANSEN essentiality-evidence classes are calculated and how they should be interpreted.

Open Essentiality Page

Overview

HANSEN essentiality prediction uses ProteomeLM-Ess, a supervised cross-species predictor. A whole-proteome contextual model is trained on experimental Mycobacterium tuberculosis essentiality and transferred to M. leprae, where it assigns each protein an essentiality probability and class.

Interpret the output as an AI-derived, cross-species essentiality estimate, not as native experimental M. leprae knockout evidence. M. leprae has a reduced genome and its biology may differ from that of M. tuberculosis.

Training labels

The model is trained on published genome-wide M. tuberculosis H37Rv essentiality calls derived by saturating transposon-sequencing (Tn-seq) mutagenesis (PMID: 28096490). Essential (ES), essential-domain (ESD) and growth-defect (GD) calls are treated as essential (625 proteins); non-essential (NE) and growth-advantage (GA) calls as non-essential (3,277 proteins). Uncertain genes are excluded from training.

Features: ESM-C + ProteomeLM

Each protein in the H37Rv (3,997 proteins) and M. leprae (1,605 proteins) proteomes is encoded with ESM-C (600M) and passed through ProteomeLM-L to obtain a whole-proteome contextual embedding, so each protein is represented relative to the rest of its proteome rather than in isolation. Because ProteomeLM accepts up to 512 proteins per pass, each proteome is processed in overlapping, genome-ordered windows and the contextual embeddings averaged.

Classifier and validation

A class-balanced logistic-regression head is trained on the H37Rv contextual embeddings, with five-fold cross-validation grouped by mmseqs2 sequence-identity clusters (≥40%) so that homologues never span the training and test split. On this leak-free M. tuberculosis cross-validation, the model reaches AUROC 0.84 (AUPRC 0.54), outperforming an ESM-C-only baseline (0.81) and the previous ESM-2 transfer model (0.71). Transferred to the 241 ortholog-anchored M. leprae proteins it gives AUROC 0.78. Top-ranked predictions recover canonical essentials and validated anti-mycobacterial drug targets (rpoB, gyrB, embC, qcrB, ribosomal proteins, sigA, infB).

Transfer to M. leprae and M. tuberculosis anchor

The trained head is applied to the M. leprae contextual embeddings to produce the ProteomeLM-Ess probability. Where a confident DIAMOND reciprocal-best-hit M. tuberculosis orthologue exists, its experimental essentiality call is shown alongside the prediction as orthogonal supporting evidence (the “M. tuberculosis anchor” column).

Essentiality classes

Each M. leprae protein is assigned a class from its ProteomeLM-Ess probability:

ClassProteomeLM-Ess probabilityInterpretation
Essential≥ 0.70High-confidence predicted essential; strongest target-discovery candidates.
Likely essential0.40 – 0.70Probably essential; warrants manual review.
Uncertain0.20 – 0.40Borderline evidence; not confidently essential or non-essential.
Non-essential< 0.20Predicted non-essential. This is a model prediction, not an experimental knockout result.

Recommended use

Use the Essentiality page as a triage tool. For drug discovery, the strongest candidates are Essential or Likely essential proteins that also have strong HANSEN target-priority evidence, good model confidence and convincing pocket or AF2Bind support.

Important: these are AI-derived, cross-species predictions transferred from M. tuberculosis experimental data, not native experimental M. leprae essentiality measurements. A “Non-essential” prediction is not an experimentally proven non-essential call. High-ranking candidates require manual biological review and experimental validation.

Lep-AMR: antimicrobial resistance analysis

How to run and interpret the Lep-AMR amplicon-based antimicrobial-resistance workflow for Mycobacterium leprae.

Open Lep-AMR

Overview

Lep-AMR is a targeted antimicrobial-resistance (AMR) analysis workflow for M. leprae based on Oxford Nanopore MinION amplicon sequencing, processed through the EPI2ME Amplicon workflow (wf-amplicon). It focuses on the established resistance-determining regions of rpoB (rifampicin), folP1 (dapsone) and gyrA (fluoroquinolones), and supports variant-level interpretation of drug resistance.

What you provide

How it works

  1. Reads are quality-filtered and aligned to the M. leprae reference resistance loci.
  2. The EPI2ME wf-amplicon workflow performs variant calling across the target amplicons.
  3. Variants within the resistance-determining regions are annotated and interpreted for rifampicin, dapsone and fluoroquinolone resistance.

Interpreting results

Each run reports the variants detected at the resistance loci together with their predicted resistance association. Known resistance-conferring mutations — for example in the rpoB rifampicin resistance-determining region (RRDR), folP1, and the gyrA quinolone resistance-determining region (QRDR) — are flagged, while novel or uncharacterised variants are reported for manual review. Calls should be interpreted in the appropriate clinical and epidemiological context and confirmed where required.

Recommended use

Use Lep-AMR for rapid, locus-targeted resistance screening of M. leprae directly from amplicon sequencing data. It complements the structural and target-discovery resources in HANSEN by linking observed resistance mutations to their drug targets. Open the workflow from the Lep-AMR page.

Contact and acknowledgements

HANSEN was developed by Dr Sundeep Chaitanya Vedithi in collaboration with the Department of Biochemistry and Department of Medicine, University of Cambridge, and the Science and Technology Facilities Council (STFC).

This work was supported by Hope Rises International, formerly American Leprosy Missions. For enquiries, please visit the Contact page.