Skip to content

Latest commit

 

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

License: MIT

Bacterial Secretion Systems Identification

Abstract

This repository contains a comprehensive computational pipeline designed for the genome-wide identification, annotation, and characterization of bacterial macromolecular secretion systems, conjugation machineries, type IV filaments, and CRISPR-Cas systems. The primary reference organism analyzed in this framework is Bacteroides xylanisolvens. The underlying analytical methodology employs probabilstic Hidden Markov Model (HMM) profile searches via the MacSyFinder framework, utilizing highly specialized model databases including TXSScan, CONJScan, TFFscan, and CasFinder.

This repository contains the exact workflow, environments, and visualization scripts required to reproduce the genomic analysis.

Methodology

1. Data Acquisition

The intact proteome sequence of Bacteroides xylanisolvens was retrieved from the National Center for Biotechnology Information (NCBI) Reference Sequence database. The designated assembly accession utilized for this analysis is GCF_018289135.1 (ASM1828913v1).

2. Genetic Element Identification via MacSyFinder

The retrieved proteome was mapped and evaluated using MacSyFinder. Given that draft genomes or unordered proteomes lack conserved operon synteny inherently present in complete replicons, the pipeline explicitly invokes the unordered execution mode (--db-type unordered).

The analytical framework is modularized into discrete execution scripts, each deploying dedicated HMM profiles:

  • Secretion Systems (txsscan.sh): Employs the TXSScan module to detect and reconstruct various bacterial protein secretion systems ranging from Type I to Type IX.
  • Type IV Filaments (tffscan.sh): Deploys TFFscan models to delineate Type IV pili (T4P), tight adherence (Tad) pili, and analogous appendages.
  • Conjugative Systems (conjscan.sh): Identifies putative conjugative apparatuses using CONJScan models, essential for unravelling horizontal gene transfer potentials.
  • CRISPR-Cas Systems (casfinder.sh): Scans for CRISPR-associated (Cas) proteins denoting adaptive immune system machineries.

3. Data Processing and Visualization

Downstream processing of the resulting annotations (all_systems.tsv) is conducted via a Python-centric stack incorporating Pandas and Matplotlib. The parsing methodology applies stringent quality thresholds (e.g., system wholeness evaluation, sequence coverage $>=$ 0.8, Expectation values $<=$ 1e-10) to circumvent false positives.

The visualization modules map the absolute abundances and architectural compositions of detected nanomachinery subunits:

  • 01_plot.py: Generates a discrete quantitative distribution of individual Type I Secretion System (T1SS) structural components (i.e., OMC, ABC, and MFP).
  • 02_plot_secretion_systems.py: Formulates a multi-system compositional overview using stacked bar plots, illustrating concurrent presence and genomic integration of different secretion systems across targeted strains.

Results

Type I Secretion System (T1SS) Component Abundance

The systematic evaluation of T1SS architectures reveals distinct frequencies of essential trans-envelope elements. The chart below explicitly demonstrates the hit distributions for the outer membrane factor (omf), ABC transporter (abc), and membrane fusion protein (mfp) constituents specific to B. xylanisolvens.

T1SS Component Hits

Broad-Spectrum Secretion System Architecture

A comparative compositional analysis emphasizes the proportional genetic assembly of diverse secretion paradigms within the queried organism. By aggregating sequence analysis outputs, the distribution below outlines the functional gene repertoire spanning detected secretion modules.

Secretion Systems Barplot

Repository Structure

.
├── 01_plot.py                              # Generates T1SS component hit charts
├── 02_plot_secretion_systems.py            # Generates comprehensive stacked bar charts
├── casfinder.sh                            # Pipeline executable for CRISPR-Cas scan
├── conjscan.sh                             # Pipeline executable for Conjugation system scan
├── tffscan.sh                              # Pipeline executable for Type IV Filament scan
├── txsscan.sh                              # Pipeline executable for TXSScan execution
├── installation.sh                         # Essential runtime dependencies configuration
├── b_xylanisolvens.faa                     # Target proteomic sequence dataset
├── macsyfinder_TXSScan_b_xylanisolvens_unordered/ # Output directory for TXSScan
├── CasFinder/                              # Output directory for CasFinder scan
├── CONJScan/                               # Output directory for CONJScan
├── TFFscan/                                # Output directory for TFFscan
└── README.md                               # Project documentation

Large File Management

In compliance with standard version control paradigms and GitHub data constraints, protective measures have been emplaced to prevent committing contiguous data blocks exceeding 100 megabytes (MB). A corresponding .gitignore formulation accounts for intermediate analysis files and raw uncompressed datasets. In instances where output caches or subsequent sequence databases exceed the designated size threshold, .readme placeholders are embedded strategically to denote structural omission pending local recomputation.

Reproducibility

To ensure maximal congruence and successful replication, an isolated environment must be provisioned. Initialization requires the sequential execution of installation.sh to install baseline dependencies, followed by the specific structural detection scripts in an independent manner. Resulting matrices are dynamically imported during the plot rendering processes.

License

This project is licensed under the MIT License - see the LICENSE file for details.

About

A computational pipeline leveraging HMM-based probabilistic modeling to systematically identify, annotate, and map bacterial macromolecular secretion systems, conjugative transfer machineries, and CRISPR-Cas arrays across Bacteroidales proteomes.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages