mtase-motif finds bacterial DNA methyltransferase (MTase) candidates in a
genome and assigns recognition motifs from annotated homologs.
The default sequence-first workflow combines:
- Prodigal gene prediction, unless a protein FASTA is supplied.
- Candidate discovery from HMMER matches to curated Pfam/optional TIGRFAMs models and qualifying whole-proteome REBASE homology. Either source can admit a candidate; an HMM hit is not required for a qualifying REBASE hit.
- MMseqs2 or BLAST+ searches of candidates against motif-labeled REBASE proteins, with family, cluster and Type I specificity evidence where available.
- Conservative motif transfer reported as
linked,ambiguous, orunresolved, with explicit provenance and confidence.
When experimental motifs are already known but their enzymes are not, the
optional motif-first reverse-linking workflow ranks genome candidates against
characterized REBASE MTases. The default evidence-weighted policy can use
separate Type III M-subunit and neighboring R-subunit evidence; the optional
motif-aware policy reports that context as annotation. The sequence-first
stages run before this optional reconciliation.
The important scientific boundary is that motif assignment is homology-based.
A genome alone usually does not contain enough information to determine an
MTase recognition motif de novo. Candidates without adequate reference or
methylation evidence are reported as unresolved; they are not given a
fabricated motif.
a. Protein-domain and REBASE evidence identify candidates, followed by
reference-guided motif assignment. b. Observed motifs optionally link back
to those candidates; only a unique linked/high result can update the primary
call. c. Optional methylation support adjusts confidence before reporting;
FIMO scanning is enabled with --list-loci. Protein tracks, GATC and ranks are
schematic examples, not experimental results.
Vector diagram
and panel guide.
PNG and SVG copies are included under docs/figures/ in the source distribution.
The Python package supports Python 3.10 or newer:
python -m pip install mtasemotif
mtase-motif --versionNative tools are installed separately. A source checkout includes a conda environment with the tools needed for the core workflow:
git clone https://github.com/lrslab/mtasemotif.git
cd mtasemotif
conda env create -f environment.yml
conda activate mtase
python -m pip install .Core runs require Prodigal when proteins are not supplied, HMMER, and either MMseqs2 or BLAST+. BLAST+ also enables family-rescue and candidate-to-candidate transfer routes when MMseqs2 handles the main search. FIMO and TIGRFAMs are optional.
Downloaded databases are not bundled in the wheel or source distribution.
The default database directory is ~/.cache/mtase-motif/db.
mtase-motif db init
mtase-motif db fetch pfam
mtase-motif db fetch rebase
mtase-motif db index
mtase-motif db statusThe default REBASE fetch uses a compact motif-labeled protein set plus REBASE
withref records for exact methylation types (for example, Dam/GATC as m6A).
It also attempts to stage role-specific expanded protein sets used only when
--known-motifs is supplied; these do not enlarge the routine whole-proteome
search. Database caches created before version 0.3.0 must fetch REBASE again
and rebuild their indexes before motif-first reverse linking can be used:
mtase-motif db fetch rebase
mtase-motif db indexThe Pfam subset in version 0.4.0 also includes the Eco57I methylase model
(PF07669). To add it to an existing database, run mtase-motif db fetch pfam
and then mtase-motif db index. Candidate thresholds remain unchanged.
This detects an Eco57I-family protein domain, not the Eco57I recognition
sequence itself; motif specificity still requires reference evidence.
Re-fetch and re-index REBASE as well to populate the reference chemistry,
modified-position and provenance fields added in 0.4.0. Older maps remain
readable, with absent metadata treated as unknown.
Local Pfam or REBASE mirrors can be imported with
--source:
mtase-motif db fetch pfam --source /path/to/Pfam-A.hmm.gz
mtase-motif db fetch rebase --source /path/to/rebase_directory
mtase-motif db indexTIGRFAMs is optional and local-source-only:
mtase-motif db fetch tigrfams --source /path/to/TIGRFAMs.hmm.gz
mtase-motif db indexSee the database setup guide for offline REBASE protein requirements and the local database layout.
Pfam, TIGRFAMs, and REBASE remain third-party resources. No downloaded records are redistributed by this package; users are responsible for their respective access, license, and citation requirements.
Check the genome, database, and native tools before starting:
mtase-motif doctor --genome genome.fnaAdd --json for a machine-readable report and inspect its top-level ok and
issues fields before launching a long run.
Run the complete sequence-first workflow:
mtase-motif run \
--genome genome.fna \
--out results/genome \
-j 4Use --db-dir with every database and run command when using a non-default
database location. Gene prediction uses translation table 11 by default.
Use --genetic-code 4 when the input genome requires table 4; this selection
is recorded in the run manifest. Choose the code from the input's biological
annotation, rather than a predicted motif. If proteins are already available,
skip Prodigal:
mtase-motif run \
--genome genome.fna \
--proteins proteins.faa \
--db-dir databases/mtase-motif \
--out results/genome \
-j 4--genome is required even when a protein FASTA is supplied. Supplying proteins
skips gene prediction, but the current interface does not accept matching
genomic coordinates. Neighborhood evidence is therefore unavailable for that
mode; prefer the genome-only route when Type I or Type III system context is
important. --genetic-code does not reinterpret supplied proteins;
the manifest records a null genetic code when gene prediction is skipped.
The default balanced motif profile is appropriate for routine use. The
sensitive profile admits more remote REBASE family evidence:
mtase-motif run \
--genome genome.fna \
--out results/genome-sensitive \
--motif-sensitivity sensitiveThe output layout separates the routine result from audit-level detail:
| File | Contents |
|---|---|
results.tsv |
Compact main result; exactly one row and 25 analysis-ready columns per MTase candidate |
details/mtase_candidates.tsv |
Candidate-discovery audit table; exactly one row per candidate |
details/motif_assignments.tsv |
Assignment evidence; one or more rows per candidate for primary, alternate, hint, or unresolved routes |
details/methylation_support.tsv |
Experimental counts and confidence effects; generated only with --methylation-support |
details/known_motif_links.tsv |
Ranked motif-to-enzyme evidence; generated only with --known-motifs |
details/motif_loci.tsv |
Motif counts and density; generated only with --list-loci |
run_manifest.json |
Schema/software versions, privacy-safe input hashes, database version, parameters, tool versions, relative outputs, and summary counts |
artifacts/<candidate_id>/motif/pwm.meme |
MEME-format PWM; empty placeholder when no motif is assigned |
candidate_id links results.tsv to candidate-oriented detail rows.
details/known_motif_links.tsv is instead organized by observed motif; it may
contain several candidate rows per motif or a blank candidate for an
unresolved motif. Its 39-column schema includes policy and reference metadata;
read detail fields by column name. results.tsv is the table to use for routine
analysis; the detail tables explain discovery evidence and every assignment route. Internal
files under work/ are deleted after a successful validation; use
--keep-work only when debugging. The CLI validates IDs, ranks, and summary
counts across the public outputs before reporting success.
The main assignment_state is linked when a primary motif is selected,
ambiguous when only hint or related-candidate evidence remains, or unresolved
when no assignment is available. This candidate-level state is separate from the motif-level
link_state in details/known_motif_links.tsv. A selected motif can have low,
medium or high confidence; linked alone does not mean experimental validation.
For methylation fields, methylation is the normalized class (m6A, m5C,
m4C, or a slash-separated combination), and mod_position is 1-based within
the displayed motif.
methylation_source and mod_position_source explain how the main call was
obtained. Raw REBASE notation such as 2(6) and conflict details remain in the
assignment audit table. Named Dam/Dcm fallback is used only when both the
enzyme name and its canonical motif agree. With no experimental support input,
the main support state is not_provided; count and fraction columns are not
added to the main table.
See the output schema for the table relationships, field definitions, controlled values, and null semantics.
When the motifs have already been measured but their enzymes are unknown, provide a motif-first TSV:
known_motif_id motif_iupac methylation mod_position source
motif_1 CGAAG m6A experiment
motif_2 CATCTC m6A 2 experimentThe input file is UTF-8 and tab-delimited, with one motif per row. The
preferred exact, lower-case header is motif_iupac, and a minimal file can
contain that single column. Compatibility aliases are documented in the full
guide. methylation is m6A, m5C, m4C, or a slash-separated combination;
mod_position is a plain 1-based integer within the displayed motif and must
be compatible with the stated chemistry. Leave either field empty when it is
unknown. An absent experimental position stays unknown after linking.
Reverse-complement input rows remain distinct events. Ready-to-run examples
are provided for
AP1
and
E. coli.
Run:
mtase-motif run \
--genome genome.fna \
--known-motifs known_motifs.tsv \
--out results/genome-known \
-j 4Choose the reverse-linking policy with --known-motif-policy:
| Behavior | evidence-weighted (default) |
motif-aware (optional) |
|---|---|---|
| Motif agreement | Exact, IUPAC-compatible, contained, or one-position-difference evidence | Identical complete motif strings in either orientation |
| Ranking | Weighted reference homology, candidate evidence, chemistry and Type III context | Bitscore, then lower e-value, query coverage, reference coverage and identity |
| Ambiguity | Near-tied weighted scores remain ambiguous | Exact numerical ties between candidates remain ambiguous |
| Context | Type III evidence contributes to the score; architecture can lower final confidence | Type III context and architecture are annotations for the accepted link |
link_score |
Heuristic score, not a probability | Blank; no composite score |
The motif-aware alignment gates are e-value ≤1e-5, identity ≥30%, and both
query and reference coverage ≥50%. --motif-sensitivity and sequence-first
threshold overrides do not change these gates. Explicit reference chemistry,
position or source conflicts exclude a reference; missing metadata stays
unknown and is reported in warnings. In this policy, linked/high means a
unique qualifying winner, not a calibrated probability or verified function.
This writes details/known_motif_links.tsv, with ranked genome candidates for
each observed motif. Under the default policy, exact or compatible motif
agreement and strong full-length protein homology provide the primary evidence.
A nearby independently detected Type III R subunit can raise a Type III M
link to high confidence. A
one-position motif difference is capped at medium confidence, close candidates
remain ambiguous, and a candidate strongly supporting another known motif is
penalized for mismatch-only explanations. Only a unique linked/high result
can replace a sequence-first primary call; medium and ambiguous links remain hints. The link
score is a ranking heuristic, not a probability. See the
known-motif linking guide
for details.
Provide a normalized support table to validate sequence-derived calls:
mtase-motif run \
--genome genome.fna \
--out results/genome \
--methylation-support methylation_support.tsvThe table requires motif_iupac, methylated_instances and
unmethylated_instances; candidate_id, methylation and mod_position are
optional. Support or contradiction can adjust confidence after sequence-first
assignment and any known-motif reconciliation, under either linking policy.
The link table retains the evidence at the linking stage; final confidence
and support effects appear in results.tsv and details/methylation_support.tsv.
Candidate-specific evidence may rescue unresolved calls only when
--methylation-rescue-unresolved is explicitly enabled. Generic motif counts
do not identify which enzyme produced that motif.
Hammerhead motif output can be converted to this schema:
mtase-motif convert-hammerhead-support \
--motifs-tsv motifs.tsv \
--out-tsv methylation_support.tsvAdd --list-loci to scan the genome for resolved motifs and write FIMO and QC
outputs plus details/motif_loci.tsv. The detail table still contains one row
per candidate; unresolved candidates receive a blank motif, zero counts, and a
no_motif warning. This requires FIMO from the MEME suite.
Successful runs remove the large internal work/ directory by default. Keep
it for troubleshooting with:
mtase-motif run \
--genome genome.fna \
--out results/genome \
--keep-workpython -m pip install -e '.[dev]'
make lint
make test
make package-checkReleases are built by GitHub Actions from tags matching the package version.
After updating mtase_motif.__version__ and CHANGELOG.md, create a tag such
as vX.Y.Z. The workflow verifies the tag, builds the wheel and source
distribution, publishes through PyPI Trusted Publishing, and then creates the
GitHub release using the matching CHANGELOG.md section as its release notes.
Use CITATION.cff for the software citation and report the package version or source commit used in your analysis. The citation currently uses the project's existing collective author, LRSLab.
The package source is released under the MIT License. Downloaded third-party database content is not covered by this license and is not bundled.
