1. What HANSEN contains
HANSEN is an integrated structural and functional resource for the Mycobacterium leprae proteome. It brings together core protein identifiers, curated functional annotations, sequence features, cross-references, homology and de novo structure models, oligomeric assemblies, ligand annotations, pocket predictions, AF2Bind binding-site predictions, PAE confidence maps and B-cell epitope propensity information.
2. Home page and search
The home page is the main entry point. Use the query panel to search by identifier, name, ligand or sequence. The top navigation bar provides direct access to the query area, database background, statistics dashboard and GO browser.
Use the query panel when you know a target
Search using a UniProt accession, ML locus tag, gene name, protein name, ligand symbol, ligand name, PubChem CID or protein sequence fragment.
Use global pages when you are exploring
Open the Statistics dashboard for database-wide coverage or the GO Browser to discover groups of proteins by ontology.
About HANSEN and database scope
This section summarises the resource scope, while the home page remains focused on search and navigation.
About HANSEN
HANSEN supports translational and mechanistic leprosy research by making M. leprae protein-centric information easy to search, inspect and reuse. The platform links identifiers, functional evidence, sequence features and structural models in a format suitable for target triage, biomarker assessment, comparative interpretation and downstream experimental design.
HANSEN scope
3. Practical example links
Use these examples to understand how HANSEN links database-wide views to protein-level interpretation. They are designed as quick entry points for new users.
4. Query types
| Query type | What to enter | What HANSEN does | Best use |
|---|---|---|---|
| UniProt | UniProt accession or entry identifier | Opens the matching protein page. | When you know the canonical UniProt entry. |
| ML locus tag | Example: ML0005 |
Finds the protein by M. leprae locus tag. | Best for genome and proteome workflows. |
| Gene name | Gene symbol or partial gene text | Uses exact, prefix and partial-text matching. | When working from annotation tables or literature. |
| Protein name | Full or partial protein description | Finds the closest matching annotated protein name. | When you know the biological function but not the locus tag. |
| Sequence | Full or partial amino-acid sequence | Normalises the sequence and searches for exact, contained and local matches. | Useful for fragments, peptides or copied FASTA sequence. |
| Ligand | Ligand code, name, SMILES-derived annotation, or PubChem CID | Returns proteins linked to that ligand through imported oligomer/template ligand annotations. | Useful for identifying ligand-associated targets or cofactor-binding proteins. |
5. Protein result page
A protein result page is the main detailed view for a single HANSEN entry. It usually begins with a summary panel containing the UniProt entry, ML locus tag, gene name, protein name, organism, length and other core identifiers. The rest of the page is divided into functional, structural and evidence sections.
ASummary and identifiers
Use this section to confirm that you have opened the correct protein. It contains entry names, gene names, ML locus tags and links to external resources where available.
BFunctional annotation
Review EC numbers, catalytic activity, cofactors, pathways, binding sites, active sites, protein family, domains, motifs and other curated annotations.
CEvidence and cross-references
Check PubMed IDs, DOI IDs, InterPro, KEGG, GeneID, STRING and other cross-references to connect HANSEN annotations with external evidence.
DStructure and target-discovery panels
Inspect model quality, Mol* structures, PAE maps, pockets, AF2Bind predictions, ligands and epitope predictions to prioritise proteins.
6. Structure models
The model section lists the structural models available for the selected protein. Depending on data availability, HANSEN can show monomeric and oligomeric models generated by AF3, Boltz, Chai and Boltz-2.
| Model type | Meaning | How to use it |
|---|---|---|
| Monomer | Single-protein structural model. | Best for inspecting domain architecture, local residues, active sites and local confidence. |
| Homomer | Oligomer made from repeated copies of the same ML protein. | Use to inspect biological assemblies, repeated interfaces and symmetry-related pockets. |
| Heteromer | Complex containing different protein components. | Use to inspect protein-protein interfaces, complex-specific pockets and chain-level behaviour. |
| Template-linked oligomer | Assembly reconstructed or modelled using template evidence. | Use template PDB and assembly metadata to interpret biological relevance. |
7. Mol* viewer
The Mol* panel is the interactive 3D structure viewer. Use it to rotate and zoom the model, inspect chains, focus on residues or ligands, and view pocket or AF2Bind overlays when available.
- Load or select a model: choose a structure model from the model controls. Oligomers may take longer to load than monomers.
- Rotate and zoom: drag to rotate, scroll to zoom and right-click or secondary-drag to pan, depending on your device.
- Use full-page mode: open the larger Mol* view when detailed inspection is needed.
- Focus features: use ligand, pocket, AF2Bind or epitope controls to focus the relevant region in 3D.
- Return to the page: exit full-page mode to continue reviewing tables and annotation cards.
8. pLDDT and PAE confidence interpretation
Confidence panels help distinguish reliable local structural regions from lower-confidence or flexible areas. pLDDT reports residue-level confidence, while PAE helps assess the reliability of domain-domain or chain-chain placement.
pLDDT
Use pLDDT to judge local residue confidence. High-confidence residues are more reliable for interpreting local pockets, active sites and epitopes. Low-confidence regions may be flexible, disordered or modelled with greater uncertainty.
PAE
Use PAE to judge relative placement of domains and chains. For oligomers, inspect block patterns across chains to understand interface confidence and assembly reliability.
9. Pockets, hotspots and AF2Bind
HANSEN includes predicted pockets and AF2Bind-style binding-site predictions where available. These panels help prioritise proteins and residues for druggability assessment.
- Pocket table: lists predicted pockets, scores, tools and residue information where available.
- Focus pocket: centres Mol* on the selected pocket and overlays the pocket region if residue-level data are available.
- AF2Bind: highlights predicted binding residues and can be used alongside pocket predictions.
- Remove/hide overlays: clear surfaces after inspection to return to a clean structure view.
For target prioritisation, favour proteins whose predicted pockets overlap with high-confidence structural regions, biologically relevant ligands, catalytic sites, conserved residues or essential functional annotations.
10. Ligands
Ligand annotations help connect modelled oligomers and template assemblies with cofactors, ions, substrates or other small molecules. Ligands may appear as CCD-style symbols, names, SMILES-derived entries or PubChem-linked compounds.
- Search by code: examples include short ligand symbols such as
ZNor template ligand codes. - Search by name: enter a ligand or compound name if the imported annotation includes it.
- Search by PubChem CID: use formats such as
CID 32051when PubChem resolution is available. - Protein page ligand table: use ligand rows to focus the ligand in Mol* and open PubChem where linked.
11. B-cell epitope propensity
The B-cell epitope section shows residue-level or region-level epitope propensity where available. Use it to identify exposed, structurally plausible antigenic regions, especially when the predictions are considered alongside pLDDT, accessibility and 3D localisation.
- Review the epitope table for predicted high-scoring residues or regions.
- Focus predicted residues in Mol* to evaluate their structural context.
- Compare epitope propensity with pockets, ligands and functional domains when prioritising diagnostic or immunological targets.
12. GO browser
The GO browser supports ontology-guided exploration. It uses a sunburst-style visualisation and associated protein tables to help users move from broad ontology categories to specific protein sets.
- Hover: inspect a GO branch and see its context.
- Mouse wheel: move inward or outward through GO hierarchy levels.
- Click: lock the selected branch or term.
- Review matched proteins: use the table below the sunburst to open individual protein pages.
- Use namespace filters: compare biological process, molecular function and cellular component annotations.
GO browsing is useful when you do not know a specific protein target but want to explore all proteins annotated with a functional process, molecular activity or localisation.
13. Statistics dashboard
The Statistics page gives a database-wide summary of structural coverage, model confidence, assembly types, oligomeric states, ligand codes, pocket-score classes and epitope evidence.
| Dashboard area | What it tells you | How to use it |
|---|---|---|
| Entry and annotation counts | How many proteins have core annotation, GO terms, EC numbers or 3D structural information. | Assess annotation completeness. |
| Model coverage | How many proteins have monomeric, homomeric or heteromeric models. | Identify modelling gaps and completed coverage. |
| pLDDT and PAE classes | Confidence distribution across models. | Prioritise high-confidence models for downstream interpretation. |
| Ligands | Most frequent ligand annotations and linked proteins. | Find cofactor-associated or ligand-template-associated proteins. |
| Pockets and epitopes | Distribution of pocket predictions, pocket scores and epitope evidence. | Shortlist targets for druggability or diagnostic antigen exploration. |
14. Example workflows
- Go to Query, or try a direct example such as ML0005, gyrA, or ZN ligand search.
- Enter an ML locus tag, for example
ML0005. - Open the protein page and confirm the summary identifiers.
- Inspect model scores and select the most relevant monomer or oligomer.
- Use Mol* to inspect residues, ligands, pockets and chain interfaces.
- Load PAE if you need domain or chain-placement confidence.
- Use the Ligand query on the home page, or open an example such as ZN.
- Enter a ligand symbol, name or PubChem CID.
- Review the list of proteins linked to that ligand.
- Open each protein and inspect the ligand table and Mol* focus action.
- Combine ligand evidence with pockets, confidence and functional annotation.
- Open the GO Browser.
- Select a namespace or navigate through the sunburst.
- Click a GO branch or term to lock the protein set.
- Open candidate proteins from the table.
- Inspect structures, pockets, ligands and epitope predictions for each candidate.
- Start with Statistics to identify proteins with models, pockets and ligand annotations.
- Open proteins with high-confidence structures and biologically relevant assemblies.
- Prioritise proteins with strong pockets, AF2Bind residues, ligands or conserved functional sites.
- Check PAE for domain/interface reliability before interpreting oligomeric pockets.
- Use external links and annotations to assess biological relevance.
15. Troubleshooting and interpretation tips
| Issue | Likely reason | What to do |
|---|---|---|
| Search gives no protein | The query may not match the stored identifier or annotation text. | Try UniProt, ML locus tag, gene name and partial protein name separately. |
| Sequence search gives no result | The fragment may be too short, absent, or from a different strain/annotation set. | Use a longer sequence fragment and remove spaces, numbers and FASTA headers. |
| Oligomer loads slowly | Large CIF/BCIF and PAE files can be heavy. | Load the structure first; load PAE only when needed. |
| PAE is missing | The PAE JSON may not exist for that selected model or may not be mapped. | Use another model or check whether the corresponding PAE file has been staged. |
| Pocket focus works but no surface appears | Residue lists may be unavailable for that pocket source/model. | Use available residue-level pocket outputs or rebuild pocket summaries with residues. |
| Ligand search misses expected ligands | Reconstructed YAML ligands may not yet be imported into the ligand table. | Ask the HANSEN maintainer to import the reconstructed ligand annotations, then refresh the page. |
Target prioritisation
The Target Prioritisation page ranks Mycobacterium leprae proteins by integrating ProteomeLM-derived proteome-context information with structural, functional and druggability evidence.
What the page shows
The table provides an evidence-weighted shortlist of candidate proteins for target discovery. Each row corresponds to an ML locus or protein and includes a Target Priority Score, priority tier, ProteomeLM contextual signal, pocket/AF2Bind evidence, model-quality evidence, annotation support and a short rationale explaining why the protein was ranked.
How the ProteomeLM signal was generated
Each protein is converted into an ESM-C (600M) embedding and passed through ProteomeLM-L in whole-proteome context, so it is interpreted relative to the rest of the M. leprae proteome rather than as an isolated sequence. A supervised ProteomeLM-Ess head, based on logistic regression of the contextual embeddings, is then trained on experimental M. tuberculosis transposon-sequencing (Tn-seq) essentiality labels (625 essential and 3,277 non-essential proteins), using mmseqs homology-grouped cross-validation. The trained model is transferred to score every M. leprae protein. In leak-free grouped cross-validation in M. tuberculosis, it reaches AUROC 0.84, compared with 0.81 for ESM-C and 0.71 for the earlier ESM-2 model. This ProteomeLM-Ess probability is now the primary ProteomeLM signal in the Target Priority Score.
Evidence layers used
| Evidence layer | Role in prioritisation |
|---|---|
| ProteomeLM contextual signal | AI-derived whole-proteome contextual evidence from ESM-C embeddings and ProteomeLM inference. |
| Boltz-2 model quality | Supports confidence in structure-based interpretation, including pockets and residue-level signals. |
| Pocket evidence | Captures predicted cavities and potential small-molecule binding sites. |
| AF2Bind evidence | Summarises residue-level binding propensity and high-propensity binding regions. |
| Functional annotation | Uses EC, GO, pathway, active-site, binding-site, cofactor and protein-family information. |
| Tractability features | Considers properties such as protein size, enzymatic function and structural druggability support. |
How the Target Priority Score is calculated
The final score is a composite 0–100 score combining AI-derived context, druggability, annotation, model quality and tractability:
Target Priority Score =
100 × (
0.35 × ProteomeLM contextual score +
0.25 × pocket/AF2Bind binding-site score +
0.20 × annotation score +
0.10 × Boltz-2 model-quality score +
0.10 × tractability score
)
Priority tiers
- High-priority: strongest combined evidence; suitable for immediate manual inspection.
- Strong candidate: good combined evidence; suitable for target shortlisting and review.
- Moderate candidate: some supportive evidence, but one or more components may be missing or weaker.
- Exploratory: lower current evidence; useful for hypothesis generation or later re-analysis.
How to use the ranking
Use the table as a discovery shortlist rather than as a final decision. The most compelling targets are those with a high Target Priority Score, strong ProteomeLM contextual signal, confident Boltz-2 model, clear pocket or AF2Bind evidence, relevant pathway or enzyme annotation and biological plausibility.
Current limitation
HANSEN now provides a supervised ProteomeLM-Ess essentiality probability, trained on M. tuberculosis transposon-sequencing (Tn-seq) essentiality labels and transferred to M. leprae through the shared ProteomeLM contextual feature space. The remaining limitation is biological: these are cross-species predictions, not native experimental M. leprae knockout data. They should therefore guide, rather than replace, experimental target validation.
HANSEN essentiality prediction
This section explains how HANSEN essentiality-evidence classes are calculated and how they should be interpreted.
Overview
HANSEN essentiality prediction uses ProteomeLM-Ess, a supervised cross-species predictor. A whole-proteome contextual model is trained on experimental Mycobacterium tuberculosis essentiality and transferred to M. leprae, where it assigns each protein an essentiality probability and class.
Interpret the output as an AI-derived, cross-species essentiality estimate, not as native experimental M. leprae knockout evidence. M. leprae has a reduced genome and its biology may differ from that of M. tuberculosis.
Training labels
The model is trained on published genome-wide M. tuberculosis H37Rv essentiality calls derived by saturating transposon-sequencing (Tn-seq) mutagenesis (PMID: 28096490). Essential (ES), essential-domain (ESD) and growth-defect (GD) calls are treated as essential (625 proteins); non-essential (NE) and growth-advantage (GA) calls as non-essential (3,277 proteins). Uncertain genes are excluded from training.
Features: ESM-C + ProteomeLM
Each protein in the H37Rv (3,997 proteins) and M. leprae (1,605 proteins) proteomes is encoded with ESM-C (600M) and passed through ProteomeLM-L to obtain a whole-proteome contextual embedding, so each protein is represented relative to the rest of its proteome rather than in isolation. Because ProteomeLM accepts up to 512 proteins per pass, each proteome is processed in overlapping, genome-ordered windows and the contextual embeddings averaged.
Classifier and validation
A class-balanced logistic-regression head is trained on the H37Rv contextual embeddings, with five-fold cross-validation grouped by mmseqs2 sequence-identity clusters (≥40%) so that homologues never span the training and test split. On this leak-free M. tuberculosis cross-validation, the model reaches AUROC 0.84 (AUPRC 0.54), outperforming an ESM-C-only baseline (0.81) and the previous ESM-2 transfer model (0.71). Transferred to the 241 ortholog-anchored M. leprae proteins it gives AUROC 0.78. Top-ranked predictions recover canonical essentials and validated anti-mycobacterial drug targets (rpoB, gyrB, embC, qcrB, ribosomal proteins, sigA, infB).
Transfer to M. leprae and M. tuberculosis anchor
The trained head is applied to the M. leprae contextual embeddings to produce the ProteomeLM-Ess probability. Where a confident DIAMOND reciprocal-best-hit M. tuberculosis orthologue exists, its experimental essentiality call is shown alongside the prediction as orthogonal supporting evidence (the “M. tuberculosis anchor” column).
Essentiality classes
Each M. leprae protein is assigned a class from its ProteomeLM-Ess probability:
| Class | ProteomeLM-Ess probability | Interpretation |
|---|---|---|
| Essential | ≥ 0.70 | High-confidence predicted essential; strongest target-discovery candidates. |
| Likely essential | 0.40 – 0.70 | Probably essential; warrants manual review. |
| Uncertain | 0.20 – 0.40 | Borderline evidence; not confidently essential or non-essential. |
| Non-essential | < 0.20 | Predicted non-essential. This is a model prediction, not an experimental knockout result. |
Recommended use
Use the Essentiality page as a triage tool. For drug discovery, the strongest candidates are Essential or Likely essential proteins that also have strong HANSEN target-priority evidence, good model confidence and convincing pocket or AF2Bind support.
Lep-AMR: antimicrobial resistance analysis
How to run and interpret the Lep-AMR amplicon-based antimicrobial-resistance workflow for Mycobacterium leprae.
Overview
Lep-AMR is a targeted antimicrobial-resistance (AMR) analysis workflow for M. leprae based on Oxford Nanopore MinION amplicon sequencing, processed through the EPI2ME Amplicon workflow (wf-amplicon). It focuses on the established resistance-determining regions of rpoB (rifampicin), folP1 (dapsone) and gyrA (fluoroquinolones), and supports variant-level interpretation of drug resistance.
What you provide
- Oxford Nanopore amplicon reads (FASTQ) covering the rpoB, folP1 and gyrA resistance loci.
- Optional sample or run metadata to label and track the analysis.
How it works
- Reads are quality-filtered and aligned to the M. leprae reference resistance loci.
- The EPI2ME
wf-ampliconworkflow performs variant calling across the target amplicons. - Variants within the resistance-determining regions are annotated and interpreted for rifampicin, dapsone and fluoroquinolone resistance.
Interpreting results
Each run reports the variants detected at the resistance loci together with their predicted resistance association. Known resistance-conferring mutations — for example in the rpoB rifampicin resistance-determining region (RRDR), folP1, and the gyrA quinolone resistance-determining region (QRDR) — are flagged, while novel or uncharacterised variants are reported for manual review. Calls should be interpreted in the appropriate clinical and epidemiological context and confirmed where required.
Recommended use
Use Lep-AMR for rapid, locus-targeted resistance screening of M. leprae directly from amplicon sequencing data. It complements the structural and target-discovery resources in HANSEN by linking observed resistance mutations to their drug targets. Open the workflow from the Lep-AMR page.