Skip to content

Repository files navigation

mtase-motif

mtase-motif finds bacterial DNA methyltransferase (MTase) candidates in a genome and assigns recognition motifs from annotated homologs.

The default sequence-first workflow combines:

  1. Prodigal gene prediction, unless a protein FASTA is supplied.
  2. Candidate discovery from HMMER matches to curated Pfam/optional TIGRFAMs models and qualifying whole-proteome REBASE homology. Either source can admit a candidate; an HMM hit is not required for a qualifying REBASE hit.
  3. MMseqs2 or BLAST+ searches of candidates against motif-labeled REBASE proteins, with family, cluster and Type I specificity evidence where available.
  4. Conservative motif transfer reported as linked, ambiguous, or unresolved, with explicit provenance and confidence.

When experimental motifs are already known but their enzymes are not, the optional motif-first reverse-linking workflow ranks genome candidates against characterized REBASE MTases. The default evidence-weighted policy can use separate Type III M-subunit and neighboring R-subunit evidence; the optional motif-aware policy reports that context as annotation. The sequence-first stages run before this optional reconciliation.

The important scientific boundary is that motif assignment is homology-based. A genome alone usually does not contain enough information to determine an MTase recognition motif de novo. Candidates without adequate reference or methylation evidence are reported as unresolved; they are not given a fabricated motif.

Workflow overview

Three-panel workflow: sequence-first MTase discovery and motif assignment, optional motif-first enzyme linking, and evidence-aware reporting.

a. Protein-domain and REBASE evidence identify candidates, followed by reference-guided motif assignment. b. Observed motifs optionally link back to those candidates; only a unique linked/high result can update the primary call. c. Optional methylation support adjusts confidence before reporting; FIMO scanning is enabled with --list-loci. Protein tracks, GATC and ranks are schematic examples, not experimental results.

Vector diagram and panel guide. PNG and SVG copies are included under docs/figures/ in the source distribution.

Installation

The Python package supports Python 3.10 or newer:

python -m pip install mtasemotif
mtase-motif --version

Native tools are installed separately. A source checkout includes a conda environment with the tools needed for the core workflow:

git clone https://github.com/lrslab/mtasemotif.git
cd mtasemotif
conda env create -f environment.yml
conda activate mtase
python -m pip install .

Core runs require Prodigal when proteins are not supplied, HMMER, and either MMseqs2 or BLAST+. BLAST+ also enables family-rescue and candidate-to-candidate transfer routes when MMseqs2 handles the main search. FIMO and TIGRFAMs are optional.

Database setup

Downloaded databases are not bundled in the wheel or source distribution. The default database directory is ~/.cache/mtase-motif/db.

mtase-motif db init
mtase-motif db fetch pfam
mtase-motif db fetch rebase
mtase-motif db index
mtase-motif db status

The default REBASE fetch uses a compact motif-labeled protein set plus REBASE withref records for exact methylation types (for example, Dam/GATC as m6A). It also attempts to stage role-specific expanded protein sets used only when --known-motifs is supplied; these do not enlarge the routine whole-proteome search. Database caches created before version 0.3.0 must fetch REBASE again and rebuild their indexes before motif-first reverse linking can be used:

mtase-motif db fetch rebase
mtase-motif db index

The Pfam subset in version 0.4.0 also includes the Eco57I methylase model (PF07669). To add it to an existing database, run mtase-motif db fetch pfam and then mtase-motif db index. Candidate thresholds remain unchanged. This detects an Eco57I-family protein domain, not the Eco57I recognition sequence itself; motif specificity still requires reference evidence. Re-fetch and re-index REBASE as well to populate the reference chemistry, modified-position and provenance fields added in 0.4.0. Older maps remain readable, with absent metadata treated as unknown.

Local Pfam or REBASE mirrors can be imported with --source:

mtase-motif db fetch pfam --source /path/to/Pfam-A.hmm.gz
mtase-motif db fetch rebase --source /path/to/rebase_directory
mtase-motif db index

TIGRFAMs is optional and local-source-only:

mtase-motif db fetch tigrfams --source /path/to/TIGRFAMs.hmm.gz
mtase-motif db index

See the database setup guide for offline REBASE protein requirements and the local database layout.

Pfam, TIGRFAMs, and REBASE remain third-party resources. No downloaded records are redistributed by this package; users are responsible for their respective access, license, and citation requirements.

Quick start

Check the genome, database, and native tools before starting:

mtase-motif doctor --genome genome.fna

Add --json for a machine-readable report and inspect its top-level ok and issues fields before launching a long run.

Run the complete sequence-first workflow:

mtase-motif run \
  --genome genome.fna \
  --out results/genome \
  -j 4

Use --db-dir with every database and run command when using a non-default database location. Gene prediction uses translation table 11 by default. Use --genetic-code 4 when the input genome requires table 4; this selection is recorded in the run manifest. Choose the code from the input's biological annotation, rather than a predicted motif. If proteins are already available, skip Prodigal:

mtase-motif run \
  --genome genome.fna \
  --proteins proteins.faa \
  --db-dir databases/mtase-motif \
  --out results/genome \
  -j 4

--genome is required even when a protein FASTA is supplied. Supplying proteins skips gene prediction, but the current interface does not accept matching genomic coordinates. Neighborhood evidence is therefore unavailable for that mode; prefer the genome-only route when Type I or Type III system context is important. --genetic-code does not reinterpret supplied proteins; the manifest records a null genetic code when gene prediction is skipped.

The default balanced motif profile is appropriate for routine use. The sensitive profile admits more remote REBASE family evidence:

mtase-motif run \
  --genome genome.fna \
  --out results/genome-sensitive \
  --motif-sensitivity sensitive

Outputs

The output layout separates the routine result from audit-level detail:

File Contents
results.tsv Compact main result; exactly one row and 25 analysis-ready columns per MTase candidate
details/mtase_candidates.tsv Candidate-discovery audit table; exactly one row per candidate
details/motif_assignments.tsv Assignment evidence; one or more rows per candidate for primary, alternate, hint, or unresolved routes
details/methylation_support.tsv Experimental counts and confidence effects; generated only with --methylation-support
details/known_motif_links.tsv Ranked motif-to-enzyme evidence; generated only with --known-motifs
details/motif_loci.tsv Motif counts and density; generated only with --list-loci
run_manifest.json Schema/software versions, privacy-safe input hashes, database version, parameters, tool versions, relative outputs, and summary counts
artifacts/<candidate_id>/motif/pwm.meme MEME-format PWM; empty placeholder when no motif is assigned

candidate_id links results.tsv to candidate-oriented detail rows. details/known_motif_links.tsv is instead organized by observed motif; it may contain several candidate rows per motif or a blank candidate for an unresolved motif. Its 39-column schema includes policy and reference metadata; read detail fields by column name. results.tsv is the table to use for routine analysis; the detail tables explain discovery evidence and every assignment route. Internal files under work/ are deleted after a successful validation; use --keep-work only when debugging. The CLI validates IDs, ranks, and summary counts across the public outputs before reporting success.

The main assignment_state is linked when a primary motif is selected, ambiguous when only hint or related-candidate evidence remains, or unresolved when no assignment is available. This candidate-level state is separate from the motif-level link_state in details/known_motif_links.tsv. A selected motif can have low, medium or high confidence; linked alone does not mean experimental validation.

For methylation fields, methylation is the normalized class (m6A, m5C, m4C, or a slash-separated combination), and mod_position is 1-based within the displayed motif. methylation_source and mod_position_source explain how the main call was obtained. Raw REBASE notation such as 2(6) and conflict details remain in the assignment audit table. Named Dam/Dcm fallback is used only when both the enzyme name and its canonical motif agree. With no experimental support input, the main support state is not_provided; count and fraction columns are not added to the main table.

See the output schema for the table relationships, field definitions, controlled values, and null semantics.

Optional evidence

Motif-first reverse linking

When the motifs have already been measured but their enzymes are unknown, provide a motif-first TSV:

known_motif_id	motif_iupac	methylation	mod_position	source
motif_1	CGAAG	m6A		experiment
motif_2	CATCTC	m6A	2	experiment

The input file is UTF-8 and tab-delimited, with one motif per row. The preferred exact, lower-case header is motif_iupac, and a minimal file can contain that single column. Compatibility aliases are documented in the full guide. methylation is m6A, m5C, m4C, or a slash-separated combination; mod_position is a plain 1-based integer within the displayed motif and must be compatible with the stated chemistry. Leave either field empty when it is unknown. An absent experimental position stays unknown after linking. Reverse-complement input rows remain distinct events. Ready-to-run examples are provided for AP1 and E. coli. Run:

mtase-motif run \
  --genome genome.fna \
  --known-motifs known_motifs.tsv \
  --out results/genome-known \
  -j 4

Choose the reverse-linking policy with --known-motif-policy:

Behavior evidence-weighted (default) motif-aware (optional)
Motif agreement Exact, IUPAC-compatible, contained, or one-position-difference evidence Identical complete motif strings in either orientation
Ranking Weighted reference homology, candidate evidence, chemistry and Type III context Bitscore, then lower e-value, query coverage, reference coverage and identity
Ambiguity Near-tied weighted scores remain ambiguous Exact numerical ties between candidates remain ambiguous
Context Type III evidence contributes to the score; architecture can lower final confidence Type III context and architecture are annotations for the accepted link
link_score Heuristic score, not a probability Blank; no composite score

The motif-aware alignment gates are e-value ≤1e-5, identity ≥30%, and both query and reference coverage ≥50%. --motif-sensitivity and sequence-first threshold overrides do not change these gates. Explicit reference chemistry, position or source conflicts exclude a reference; missing metadata stays unknown and is reported in warnings. In this policy, linked/high means a unique qualifying winner, not a calibrated probability or verified function.

This writes details/known_motif_links.tsv, with ranked genome candidates for each observed motif. Under the default policy, exact or compatible motif agreement and strong full-length protein homology provide the primary evidence. A nearby independently detected Type III R subunit can raise a Type III M link to high confidence. A one-position motif difference is capped at medium confidence, close candidates remain ambiguous, and a candidate strongly supporting another known motif is penalized for mismatch-only explanations. Only a unique linked/high result can replace a sequence-first primary call; medium and ambiguous links remain hints. The link score is a ranking heuristic, not a probability. See the known-motif linking guide for details.

Methylation support

Provide a normalized support table to validate sequence-derived calls:

mtase-motif run \
  --genome genome.fna \
  --out results/genome \
  --methylation-support methylation_support.tsv

The table requires motif_iupac, methylated_instances and unmethylated_instances; candidate_id, methylation and mod_position are optional. Support or contradiction can adjust confidence after sequence-first assignment and any known-motif reconciliation, under either linking policy. The link table retains the evidence at the linking stage; final confidence and support effects appear in results.tsv and details/methylation_support.tsv. Candidate-specific evidence may rescue unresolved calls only when --methylation-rescue-unresolved is explicitly enabled. Generic motif counts do not identify which enzyme produced that motif.

Hammerhead motif output can be converted to this schema:

mtase-motif convert-hammerhead-support \
  --motifs-tsv motifs.tsv \
  --out-tsv methylation_support.tsv

Motif loci

Add --list-loci to scan the genome for resolved motifs and write FIMO and QC outputs plus details/motif_loci.tsv. The detail table still contains one row per candidate; unresolved candidates receive a blank motif, zero counts, and a no_motif warning. This requires FIMO from the MEME suite.

Retaining intermediates

Successful runs remove the large internal work/ directory by default. Keep it for troubleshooting with:

mtase-motif run \
  --genome genome.fna \
  --out results/genome \
  --keep-work

Development and release checks

python -m pip install -e '.[dev]'
make lint
make test
make package-check

Releases are built by GitHub Actions from tags matching the package version. After updating mtase_motif.__version__ and CHANGELOG.md, create a tag such as vX.Y.Z. The workflow verifies the tag, builds the wheel and source distribution, publishes through PyPI Trusted Publishing, and then creates the GitHub release using the matching CHANGELOG.md section as its release notes.

Citation

Use CITATION.cff for the software citation and report the package version or source commit used in your analysis. The citation currently uses the project's existing collective author, LRSLab.

License

The package source is released under the MIT License. Downloaded third-party database content is not covered by this license and is not bundled.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages