This repository contains a comprehensive computational pipeline designed for the genome-wide identification, annotation, and characterization of bacterial macromolecular secretion systems, conjugation machineries, type IV filaments, and CRISPR-Cas systems. The primary reference organism analyzed in this framework is Bacteroides xylanisolvens. The underlying analytical methodology employs probabilstic Hidden Markov Model (HMM) profile searches via the MacSyFinder framework, utilizing highly specialized model databases including TXSScan, CONJScan, TFFscan, and CasFinder.
This repository contains the exact workflow, environments, and visualization scripts required to reproduce the genomic analysis.
The intact proteome sequence of Bacteroides xylanisolvens was retrieved from the National Center for Biotechnology Information (NCBI) Reference Sequence database. The designated assembly accession utilized for this analysis is GCF_018289135.1 (ASM1828913v1).
The retrieved proteome was mapped and evaluated using MacSyFinder. Given that draft genomes or unordered proteomes lack conserved operon synteny inherently present in complete replicons, the pipeline explicitly invokes the unordered execution mode (--db-type unordered).
The analytical framework is modularized into discrete execution scripts, each deploying dedicated HMM profiles:
- Secretion Systems (
txsscan.sh): Employs the TXSScan module to detect and reconstruct various bacterial protein secretion systems ranging from Type I to Type IX. - Type IV Filaments (
tffscan.sh): Deploys TFFscan models to delineate Type IV pili (T4P), tight adherence (Tad) pili, and analogous appendages. - Conjugative Systems (
conjscan.sh): Identifies putative conjugative apparatuses using CONJScan models, essential for unravelling horizontal gene transfer potentials. - CRISPR-Cas Systems (
casfinder.sh): Scans for CRISPR-associated (Cas) proteins denoting adaptive immune system machineries.
Downstream processing of the resulting annotations (all_systems.tsv) is conducted via a Python-centric stack incorporating Pandas and Matplotlib. The parsing methodology applies stringent quality thresholds (e.g., system wholeness evaluation, sequence coverage
The visualization modules map the absolute abundances and architectural compositions of detected nanomachinery subunits:
01_plot.py: Generates a discrete quantitative distribution of individual Type I Secretion System (T1SS) structural components (i.e., OMC, ABC, and MFP).02_plot_secretion_systems.py: Formulates a multi-system compositional overview using stacked bar plots, illustrating concurrent presence and genomic integration of different secretion systems across targeted strains.
The systematic evaluation of T1SS architectures reveals distinct frequencies of essential trans-envelope elements. The chart below explicitly demonstrates the hit distributions for the outer membrane factor (omf), ABC transporter (abc), and membrane fusion protein (mfp) constituents specific to B. xylanisolvens.
A comparative compositional analysis emphasizes the proportional genetic assembly of diverse secretion paradigms within the queried organism. By aggregating sequence analysis outputs, the distribution below outlines the functional gene repertoire spanning detected secretion modules.
.
├── 01_plot.py # Generates T1SS component hit charts
├── 02_plot_secretion_systems.py # Generates comprehensive stacked bar charts
├── casfinder.sh # Pipeline executable for CRISPR-Cas scan
├── conjscan.sh # Pipeline executable for Conjugation system scan
├── tffscan.sh # Pipeline executable for Type IV Filament scan
├── txsscan.sh # Pipeline executable for TXSScan execution
├── installation.sh # Essential runtime dependencies configuration
├── b_xylanisolvens.faa # Target proteomic sequence dataset
├── macsyfinder_TXSScan_b_xylanisolvens_unordered/ # Output directory for TXSScan
├── CasFinder/ # Output directory for CasFinder scan
├── CONJScan/ # Output directory for CONJScan
├── TFFscan/ # Output directory for TFFscan
└── README.md # Project documentation
In compliance with standard version control paradigms and GitHub data constraints, protective measures have been emplaced to prevent committing contiguous data blocks exceeding 100 megabytes (MB). A corresponding .gitignore formulation accounts for intermediate analysis files and raw uncompressed datasets. In instances where output caches or subsequent sequence databases exceed the designated size threshold, .readme placeholders are embedded strategically to denote structural omission pending local recomputation.
To ensure maximal congruence and successful replication, an isolated environment must be provisioned. Initialization requires the sequential execution of installation.sh to install baseline dependencies, followed by the specific structural detection scripts in an independent manner. Resulting matrices are dynamically imported during the plot rendering processes.
This project is licensed under the MIT License - see the LICENSE file for details.

