diff --git a/.github/workflows/wheels.yml b/.github/workflows/wheels.yml index 63d8d57..6c3271d 100644 --- a/.github/workflows/wheels.yml +++ b/.github/workflows/wheels.yml @@ -45,7 +45,7 @@ jobs: strategy: fail-fast: false matrix: - os: [ubuntu-latest, macos-15-intel, macos-15] + os: [ubuntu-latest, macos-13, macos-14] steps: - uses: actions/checkout@v4 @@ -81,7 +81,7 @@ jobs: CIBW_ARCHS_MACOS: auto universal2 SYSTEM_VERSION_COMPAT: 0 # Restrict builds to CPython 3.10 to 3.12 - CIBW_BUILD: "cp310-* cp311-* cp312-* cp313-* cp314-*" + CIBW_BUILD: "cp310-* cp311-* cp312-*" # Skip PyPy and musllinux builds CIBW_SKIP: "*-musllinux* pp31*-macosx* pp*" diff --git a/.readthedocs.yaml b/.readthedocs.yaml index 49867a5..fce1e1a 100644 --- a/.readthedocs.yaml +++ b/.readthedocs.yaml @@ -4,20 +4,12 @@ build: os: "ubuntu-22.04" tools: python: "3.10" - apt_packages: - - libblas-dev - - liblapack-dev - - gfortran # We recommend specifying your dependencies to enable reproducible builds: # https://docs.readthedocs.io/en/stable/guides/reproducible-builds.html python: install: - requirements: docs/requirements.txt -# - method: pip -# path: . -# extra_requirements: -# - docsbuild sphinx: configuration: docs/source/conf.py diff --git a/README.md b/README.md index 0c7839a..2cfdea2 100644 --- a/README.md +++ b/README.md @@ -15,22 +15,15 @@ [![Downloads](https://static.pepy.tech/badge/circe-py/month)](https://pepy.tech/project/circe-py) ## Description -This repository contains a Python package for inferring **co-accessibility networks from single-cell ATAC-seq data**, using [skggm](https://www.github.com/skggm/skggm) for the graphical lasso and [scanpy](https://www.github.com/theislab/scanpy) for data processing. - -You can check our preprint here for more details! 😊
-https://doi.org/10.1101/2025.09.23.678054 - -While updating the pre-processing, CIRCE's algorithm is based on the pipeline and hypotheses presented in the manuscript "Cicero Predicts cis-Regulatory DNA Interactions from Single-Cell Chromatin Accessibility Data" by Pliner et al. (2018). This original R package [Cicero](https://cole-trapnell-lab.github.io/cicero-release/) is available [here](https://www.github.com/cole-trapnell-lab/cicero-release). - - - +This repo contains a python package for inferring **co-accessibility networks from single-cell ATAC-seq data**, using [skggm](https://www.github.com/skggm/skggm) for the graphical lasso and [scanpy](https://www.github.com/theislab/scanpy) for data processing. +It is based on the pipeline and hypotheses presented in the manuscript "Cicero Predicts cis-Regulatory DNA Interactions from Single-Cell Chromatin Accessibility Data" by Pliner et al. (2018). This R package [Cicero](https://cole-trapnell-lab.github.io/cicero-release/) is available [here](https://www.github.com/cole-trapnell-lab/cicero-release). ## Installation The package can be installed using pip: ``` pip install circe-py ``` - and from GitHub + and from github ``` pip install "git+https://github.com/cantinilab/circe.git" ``` @@ -73,15 +66,13 @@ ci.draw.plot_connections_genes( ``` -## Usage -You can go check out our documentation for more examples!
https://circe.readthedocs.io/
-The documentation is still in building, so don't hesitate to open any issues or requests you might have in this repo. 😊 +## Comparison to Cicero R package +
Metacalls computation might create differences, but scores will be identical applied to the same metacalls (cf comparison plots below). It should run significantly faster than Cicero _(e.g.: running time of 5 sec instead of 17 min for the dataset 2)_. -## Benchmark & Comparison to Cicero R package -
Metacalls computation might create differences, but scores will be identical when applied to the same metacalls (cf comparison plots below). It should run significantly faster than Cicero _(e.g., running time of 5 sec instead of 17 min for the dataset 2)_. -
*On the same metacells obtained from the Cicero code.* +_If you have any suggestion, don't hesitate ! This package is still a work in progress :)_ +
*On the same metacells obtained from Cicero code.* -All tests run in the preprint can be found in the [circe benchmark repo](https://github.com/cantinilab/Circe_reproducibility). +All tests can be found in the [circe benchmark repo](https://github.com/cantinilab/Circe_reproducibility) ### Real dataset 2 - subsample of 10x PBMC (2021) - Pearson correlation coefficient: 0.999958 @@ -94,10 +85,12 @@ Performance on real dataset 2: ### Coming: + - Gene activity calculation -## Citation -> Trimbour R., Saez Rodriguez J., Cantini L. (2025). CIRCE: a scalable Python package to predict cis-regulatory DNA interactions from single-cell chromatin accessibility data. -bioRxiv, 2025.09.23.678054, doi: https://doi.org/10.1101/2025.09.23.678054 +## Usage +It is currently developped to work with AnnData objects. Check [this notebook](https://circe.readthedocs.io/en/latest/examples/2_Detailed_example.html) for a simple usage example. +## Citation +Trimbour Rémi (2025). Circe: Co-accessibility network from ATAC-seq data in python (based on Cicero package). Package version 0.3.6. diff --git a/docs/requirements.txt b/docs/requirements.txt index 4f1a676..9904c69 100644 --- a/docs/requirements.txt +++ b/docs/requirements.txt @@ -5,7 +5,3 @@ sphinx-autodoc-typehints sphinx-copybutton myst_nb furo -tqdm -pybiomart -circe-py -rich diff --git a/docs/source/API/ccans_modules.rst b/docs/source/API/ccans_modules.rst deleted file mode 100644 index c9fa365..0000000 --- a/docs/source/API/ccans_modules.rst +++ /dev/null @@ -1,7 +0,0 @@ -CCAN modules detection -====================== - -.. automodule:: circe.ccan_module - :members: - :undoc-members: - :show-inheritance: \ No newline at end of file diff --git a/docs/source/API/download.rst b/docs/source/API/download.rst deleted file mode 100644 index 7ff3084..0000000 --- a/docs/source/API/download.rst +++ /dev/null @@ -1,7 +0,0 @@ -Download gene coordinates -=============================================== - -.. automodule:: circe.downloads - :members: - :undoc-members: - :show-inheritance: \ No newline at end of file diff --git a/docs/source/API/draw.rst b/docs/source/API/draw.rst deleted file mode 100644 index 8b88f35..0000000 --- a/docs/source/API/draw.rst +++ /dev/null @@ -1,7 +0,0 @@ -Drawing Utilities -================= - -.. automodule:: circe.draw - :members: - :undoc-members: - :show-inheritance: \ No newline at end of file diff --git a/docs/source/API/infer_network.rst b/docs/source/API/infer_network.rst deleted file mode 100644 index 411c9c4..0000000 --- a/docs/source/API/infer_network.rst +++ /dev/null @@ -1,7 +0,0 @@ -Infer network -================= - -.. automodule:: circe.circe - :members: - :undoc-members: - :show-inheritance: \ No newline at end of file diff --git a/docs/source/API/metacells.rst b/docs/source/API/metacells.rst deleted file mode 100644 index 07b508b..0000000 --- a/docs/source/API/metacells.rst +++ /dev/null @@ -1,7 +0,0 @@ -Metacells -========= - -.. automodule:: circe.metacells - :members: - :undoc-members: - :show-inheritance: \ No newline at end of file diff --git a/docs/source/circe_explained/1_circe_installation.rst b/docs/source/circe_explained/1_circe_installation.rst index dbfdd3c..0a092a8 100644 --- a/docs/source/circe_explained/1_circe_installation.rst +++ b/docs/source/circe_explained/1_circe_installation.rst @@ -22,18 +22,18 @@ Essential dependencies Since CIRCE depends on `lapack `__ and `blas `__ to compute the graphical lasso, you may need to install these libraries first. It means that you might not be able to use CIRCE in a standard Windows environment without WSL or similar. -For that, you can use a singularity or docker container, such as provided here (`singularity `__). +For that, you can use a singularity or docker container, such as provided here (`singularity `__). -If you are working in a conda environment and are encountering any issue with the installation of circe-py, you can try installing first these dependencies via: +If you are working in a conda environment and are encountering any issue with the installation you can try installing these dependencies via: .. code-block:: bash conda install conda-forge::lapack - conda install conda-forge::libfortran=3 - conda install -c conda-forge cmake llvmlite=0.46 + conda install libfortran=3 + conda install gxx_linux-64 -If you are not working with conda, +IF you are not workign with conda, on Ubuntu, you can do this via: .. code-block:: bash diff --git a/docs/source/circe_explained/circe_preprocessing.md b/docs/source/circe_explained/circe_preprocessing.md index 3d6645e..5f7ebd3 100644 --- a/docs/source/circe_explained/circe_preprocessing.md +++ b/docs/source/circe_explained/circe_preprocessing.md @@ -1,76 +1,22 @@ -# 🔧 Extended Preprocessing Guidance for CIRCE -Here are some tips and observations from our benchmarks on preprocessing choices for CIRCE ([Trimbour et al., 2025](https://www.biorxiv.org/content/10.1101/2025.09.23.678054v2)). +Preprocessing choices +================================= +Here are some tips and observations from our benchmarks on preprocessing choices for CIRCE [Trimbour et al., 2025](paper_link). -## 1. Count treatment: raw counts, binarization, count correction +Single-cell vs metacells +------------------------ +- **Single-cell inputs** typically yield **better validation** versus PC-HiC interactions. +- **Metacells** can reduce compute and sparsity but showed a small AUC drop in the tested defaults. -Use **raw counts** (peak-by-cell matrix) by default. This preserves quantitative information about accessibility that seems to help CIRCE’s co-accessibility inference. +Binarization & count correction +------------------------------- +- Simple **binarization** of counts did not notably improve predictions over raw counts. +- The traditional **Cicero count correction** decreased performance in benchmarks. -**Binarization can be tried** — especially if you want to reduce biases from highly accessible peaks or differences in coverage — but be aware that in the CIRCE context it did not yield benefit. +Neighbors space +--------------- +- Build neighbors in **LSI space** rather than low-dimensional UMAP/t-SNE to preserve distances. -**Avoid applying Cicero’s “count correction”** (or other heavy total-count normalization) prior to co-accessibility inference with CIRCE, unless you have a very strong reason and you benchmark performance (because in the published report it degraded results). - - -**Additional nuance / recommendation:** -Before analysis your dataset, consider doing quality filtering and cell QC (remove low-quality cells or potential doublets), which is standard in scATAC-seq preprocessing workflows. Several scATAC best-practice resources recommend filtering based on metrics such as fragment counts, fraction of reads in peaks, TSS enrichment, etc. -[You can check guidelines here](https://www.sc-best-practices.org/chromatin_accessibility/introduction.html) - -### Notes on Cicero-Style Preprocessing and Why We Avoid It - -Cicero’s recommended preprocessing pipeline historically includes **binarization** and a **normalization step** based on total accessibility counts per cell. Our evaluations, supported by published analyses, show that neither step is beneficial for CIRCE. - -#### Binarization -Previous studies demonstrate that binarizing scATAC-seq counts does **not** improve data quality, statistical fit, or downstream biological interpretation. -- **Martens et al., 2023** - *“Here we show that binarization is an unnecessary step that neither improves goodness of fit, clustering, cell type identification nor batch integration.”* - DOI: https://doi.org/10.1038/s41592-023-02112-6 - -This aligns with our benchmarks: **binarized counts did not outperform raw counts** for co-accessibility inference. - -#### Total-count normalization -Cicero-style correction also divides accessibility by a per-cell total count. However, this assumption — that every cell should have the same total accessibility — is unsupported biologically and methodologically. - -- **Kwok et al., 2025** - *“Dividing by total count is a sound strategy for bulk sequencing [...]. However, in scATAC-seq data [...] the total count of each cell is different. Therefore, after TF (Ed.: Term Frequency transformation) transformation, the largest variation between cells will naturally be due to their denominators, that is, the total counts per cell or sequencing depth.”* - DOI: https://doi.org/10.1186/s13059-025-03735-y - - -## 2. Input type: single-cell vs metacells / pseudobulk - -**Single-cell inputs:** As already noted, single-cell resolution yielded better validation against promoter-capture Hi-C (PC-HiC) interactions in the benchmarks with CIRCE. ([Trimbour et al., 2025](https://www.biorxiv.org/content/10.1101/2025.09.23.678054v2)) - -**Metacells:** Aggregating cells (into metacells) can reduce computational cost and mitigate sparsity, but you still need to generate enough of them or it may lead to decreased AUC. - -**When to use metacells:** -
a. If the dataset is extremely large (many tens/hundreds of thousands of cells) and memory/compute is limiting — or if individual cells are very sparse — metacells may be acceptable. But you should re-evaluate performance (e.g., link recovery, network robustness) compared to a single-cell run. -
b. If you need both **negative and positive co-accessibility scores**, and that you observe a **very high proportion of positive score** and very few negative co-accessibility scores. It typically indicate that you have a very high level of sparsity and metacells helps correcting this bias. - -#### Metacell algorithms: Dimensionality reduction & neighbor graph construction (“neighbors space”) - -To identify neighboring cells (i.e. define cell proximity / similarity), using LSI (latent semantic indexing) space rather than low-dimensional nonlinear embeddings such as UMAP or t-SNE works better to preserve relative distances. - -Rationale: Nonlinear embeddings (UMAP/t-SNE) are optimized for visualization and can distort distances in a way that may not reflect true similarity structure relevant for co-accessibility. LSI (or other linear/distance-preserving dimensionality reduction) tends to be more robust for neighbor-graph building. - -Recommendation: Use LSI (or PCA / other linear methods) for constructing the neighbor graph. Avoid using UMAP or t-SNE for that purpose. - -Tip / extension: Depending on dataset size and complexity, you might consider exploring different numbers of LSI components (e.g. carry out a small sweep: 50, 100, 200 LSI dims) to see how stable the inferred networks / CCANs are. This can help choose a setting robust to technical noise and overfitting. - -## 3. AnnData matrix format choice - -CIRCE can handle both sparse and dense matrices. A dense matrix will be **faster to process**, since the graphical lasso model ultimately needs dense matrix chunks. -On large atlases, you should store your peak-by-cell matrix in **CSC format** ([Compressed Sparse Column](https://docs.scipy.org/doc/scipy/reference/generated/scipy.sparse.csc_matrix.html)) to accelerate column extraction (i.e. per-cell operations) in CIRCE. - -## 4. Benchmarking new metacells/preprocessing strategies - -Our benchmark was limited to a comparison of Cicero and CIRCE's standard practices. - -If you want to test your own preprocessing or metacells strategy, you can have a look at our [benchmark pipeline](https://www.github.com/cantinilab/circe_reproducibility) that we used to compare the different methods. :) - -You can then simply add the Snakemake rule corresponding to your own method, and compare it to the other methods and the PC-HiC data considered there as the ground truth. -_Don't hesitate to open a GitHub issue there if you need additional guidance._ - -## Summary -- Binarization: **no demonstrated benefit** for scATAC or CIRCE. -- Total-count normalization: **methodologically unsound** for sparse single-cell chromatin data. -- Recommended: **use raw counts, without Cicero-style correction**. -- Metacells: Only if you need to correct extra-sparse data \ No newline at end of file +Matrix format tips +------------------ +- For very large atlases, store your peak-by-cell matrix as **CSC** to accelerate column extraction. diff --git a/docs/source/circe_explained/circe_scores.md b/docs/source/circe_explained/circe_scores.md index 9475bba..afec206 100644 --- a/docs/source/circe_explained/circe_scores.md +++ b/docs/source/circe_explained/circe_scores.md @@ -1,81 +1 @@ -# Understanding CIRCE Co-accessibility Scores - -CIRCE computes a co-accessibility score for each pair of genomic regions, summarizing how strongly their accessibility varies together across single cells. These scores help identify regulatory relationships, shared chromatin contexts, and coordinated patterns of chromatin opening and closing. - -Co-accessibility scores generally range from **–1 to +1**, where the sign and magnitude reflect how consistently the two regions share accessibility states. - ---- - -## Positive scores - -A **positive co-accessibility score** indicates that two regions tend to be **accessible in the same cells** and **inaccessible in the same cells**. - -Biologically, this pattern can arise from: - -- **Shared regulatory activity** - Distal enhancers, promoters, or other regulatory elements may be active in the same cell populations and therefore open together. This often reflects coordinated transcription factor binding or shared participation in a regulatory program. - -- **Chromatin architecture** - Regions within the same chromatin loop, sub-domain, or regulatory neighborhood frequently adopt similar accessibility states across cells. A positive score is a statistical signature of being in a shared 3D context, even though it does not directly measure chromatin contacts. - -- **Cell-state–dependent activation** - During differentiation or stimulus responses, groups of regulatory elements can turn on together. Positive scores highlight these co-activated modules. - -### What positive scores *do not* imply - -A positive score does **not** mean that the two regions have identical biological roles. Co-accessible elements: - -- do **not** necessarily regulate the same gene, -- may have **unequal** or unrelated regulatory impact, -- may use different transcription factors, -- do **not** guarantee physical interaction. - -Thus, co-accessibility reflects **coordinated behavior across cells**, not equivalence of function. - ---- - -## Negative scores - -A **negative co-accessibility score** means that when one region is accessible in a given cell, the other tends to be inaccessible. This can arise from: - -- **Mutually exclusive regulatory programs** - For example, two enhancers used in distinct lineages or cell states. - -- **Switch-like behavior** - Elements that participate in alternative regulatory modes during differentiation or branching trajectories. - -Negative scores generally occur less frequently and can be harder to interpret; they do not imply repression or antagonism, only **opposite accessibility patterns** across cells. - ---- - -## What co-accessibility tells you — and what it does not - -Co-accessibility scores capture the **population-level coordination** of chromatin accessibility. They are useful for generating hypotheses about regulatory architecture and for linking distal elements to possible targets. - -However, co-accessibility alone does **not** confirm: - -- direct enhancer–gene regulation, -- causal influence on gene expression, -- physical chromatin looping, -- functional equivalence between elements. - -Integrating co-accessibility with additional evidence (expression, motif enrichment, 3D contact data, etc.) provides stronger regulatory interpretation. - ---- - -## Summary - -Co-accessibility scores quantify how similarly two genomic regions behave across single cells. - -- **Positive scores** reflect coordinated accessibility and shared regulatory context. -- **Negative scores** indicate mutually exclusive or divergent accessibility patterns. -- Co-accessibility is a descriptive signal and does not, by itself, establish direct regulatory effects. - -These scores help map the structure and dynamics of the regulatory landscape, guiding downstream interpretation and hypothesis generation. - ---- -## Sources - -Pliner,H.A. et al. (2018) Cicero predicts cis-regulatory DNA interactions from single-cell chromatin accessibility data. Mol. Cell, 71, 858-871.e8. [https://doi.org/10.1016/j.molcel.2018.06.044](https://doi.org/10.1016/j.molcel.2018.06.044) - -Trimbour,R et al. (2025) CIRCE: a scalable Python package to predict cis-regulatory DNA interactions from single-cell chromatin accessibility data. bioRxiv. 2025.09.23.678054. [https://doi.org/10.1101/2025.09.23.678054](https://doi.org/10.1101/2025.09.23.678054) +# Understanding CIRCE's scores \ No newline at end of file diff --git a/docs/source/conf.py b/docs/source/conf.py index c18aed8..aea1bfa 100644 --- a/docs/source/conf.py +++ b/docs/source/conf.py @@ -20,19 +20,19 @@ import os import sys -on_rtd = os.environ.get("READTHEDOCS") == "True" -if not on_rtd: - import sys - sys.path.insert(0, os.path.abspath("../..")) - +sys.path.insert(0, os.path.abspath("../../")) + # -- General configuration --------------------------------------------------- # https://www.sphinx-doc.org/en/master/usage/configuration.html#general-configuration -extensions = ["myst_nb", "sphinx.ext.autodoc", "sphinx.ext.napoleon"] + +extensions = ["myst_nb"] jupyter_execute_notebooks = "off" templates_path = ['_templates'] exclude_patterns = [] + + # -- Options for HTML output ------------------------------------------------- # https://www.sphinx-doc.org/en/master/usage/configuration.html#options-for-html-output diff --git a/docs/source/examples/4_circe_celloracle_tutorial.ipynb b/docs/source/examples/4_circe_celloracle_tutorial.ipynb index 8bee850..dec6264 100644 --- a/docs/source/examples/4_circe_celloracle_tutorial.ipynb +++ b/docs/source/examples/4_circe_celloracle_tutorial.ipynb @@ -70,7 +70,7 @@ "id": "87d65099-2c7a-486c-88a3-072a83352400", "metadata": {}, "source": [ - "The data of this tutorial can be downloaded from [this Zenodo repository](https://zenodo.org/records/17661273)\n", + "The data of this tutorial can be downloaded at [____](www.wikipedia.org)\n", "If you have any trouble to install both CellOracle and CIRCE in the same environement, you try this conda env configuration : [CIRCE+CellOracle](https://github.com/cantinilab/Circe/blob/main/circe_celloracle_env.yml)" ] }, diff --git a/docs/source/index.rst b/docs/source/index.rst index 8784d88..ddc6185 100644 --- a/docs/source/index.rst +++ b/docs/source/index.rst @@ -1,6 +1,3 @@ -.. role:: raw-html(raw) - :format: html - .. circe documentation master file, created by sphinx-quickstart on Thu Sep 4 18:21:17 2025. You can adapt this file completely to your liking, but it should at least @@ -29,16 +26,6 @@ CIRCE: Cis-regulatory interactions between chromatin regions examples/* -.. toctree:: - :maxdepth: 1 - :hidden: - :glob: - :caption: API Reference - - API/* - - - .. image:: https://github.com/cantinilab/circe/actions/workflows/codecov.yaml/badge.svg :target: https://github.com/cantinilab/circe/actions/workflows/codecov.yaml :alt: Unit Tests @@ -64,19 +51,13 @@ CIRCE is a Python package for inferring **co-accessibility networks from single-cell ATAC-seq data**, using `skggm `_ for the graphical lasso and `scanpy `_ for data processing. -You can check our preprint here for more details! 😊 :raw-html:`
` -https://doi.org/10.1101/2025.09.23.678054 - -While updating the preprocessing, the algorithm is based on the pipeline and hypotheses presented in the manuscript +It is based on the pipeline and hypotheses presented in the manuscript *Cicero Predicts cis-Regulatory DNA Interactions from Single-Cell Chromatin Accessibility Data* by Pliner et al. (2018). The original R package Cicero is available `here `_. -.. note:: - In case you encounter any trouble, check out the `CIRCE GitHub repo `_. - Installation ------------ @@ -135,13 +116,17 @@ Visualisation :align: center -Benchmark & comparison to the Cicero R package +Comparison to Cicero R package ------------------------------ -All tests run in the preprint can be found in the `CIRCE benchmark repo `_.. -Metacells computation might cause differences, but scores will be identical when applied to the same metacells (cf. comparison plots below). -It should run significantly faster than Cicero (e.g., running time of 5 sec instead of 17 min for dataset 2). -*On the same metacells obtained from the Cicero code.* +Metacalls computation might create differences, but scores will be identical applied to the same metacalls (cf comparison plots below). +It should run significantly faster than Cicero (e.g.: running time of 5 sec instead of 17 min for dataset 2). + +If you have any suggestion, don't hesitate! This package is still a work in progress :) + +On the same metacells obtained from Cicero code. + +All tests can be found in the `circe benchmark repo `_. Real dataset 2 - subsample of 10x PBMC (2021) @@ -171,11 +156,19 @@ Coming - Complete integration in HuMMuS GRN inference pipeline +Usage +----- + +It is currently developed to work with AnnData objects. +Check [CIRCE general example](https://circe.readthedocs.io/en/latest/examples/2_Detailed_example.html) for a simple usage demonstration. + + Citation -------- -Trimbour R., Saez Rodriguez J., Cantini L. (2025). CIRCE: a scalable Python package to predict cis-regulatory DNA interactions from single-cell chromatin accessibility data. -bioRxiv, 2025.09.23.678054, doi: https://doi.org/10.1101/2025.09.23.678054 +Trimbour Rémi (2025). +*Circe: Co-accessibility network from ATAC-seq data in Python (based on Cicero package).* +Package version 0.3.6. .. toctree::