Skip to content
sukhleenkPublic

About

ArXiv digest tool that runs completely locally with summarization models.

Topics

Resources

Stars

6 stars

Watchers

2 watching

Forks

Repository files navigation

Sift

A local ArXiv research digest for macOS and Linux. Sift fetches new papers on the topics you follow, summarizes them with a HuggingFace model running on your machine, groups them by theme, and hands you a single HTML page to read.

Built in the hopes of saving time for researchers wanting to stay up-to-date with preprints.

Tests Python 3.11+ License: Apache 2.0

Sift digest page

Contents

Features

  • Local summarization. Abstracts are condensed by a transformer model on your own hardware. Offline after the first model download.
  • Topic clustering. Papers are embedded and grouped by k-means, with the number of clusters chosen by a cosine silhouette sweep instead of a fixed k. Each cluster gets a TF-IDF label.
  • Fair topic mixing. Papers are picked round-robin across your topics, so one busy topic cannot crowd out the rest.
  • No repeats. Papers that already appeared in an earlier digest are skipped.
  • Read, save, and annotate. Every paper in the digest has Mark Read, Save, and Copy BibTeX buttons. Saved papers get their own page with a notes field.
  • Digest history. Past digests stay in the menu until they age out of the retention window.
  • Scheduled or on demand. Once or twice a day at hours you choose, plus a Fetch Now item and a pause-until-tomorrow switch.
  • Hardware-aware defaults. The setup wizard detects your chip, memory, and CUDA devices and picks model sizes that fit.

Requirements

  • Python 3.11 or later
  • macOS 12+ or Linux x86_64
  • 4-16 GB free disk space, depending on the model you pick
  • Linux tray support needs GTK bindings: sudo apt install python3-gi gir1.2-ayatanaappindicator3-0.1

Installation

git clone https://github.com/sukhleenk/sift.git
cd sift
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

First run

python main.py

Running Sift from source is currently more stable than the packaged builds in dist/, and is the recommended path.

A setup wizard opens on first launch. It reports the hardware it found, suggests a summarization and embedding model, and lets you override both. You also set your topics (comma separated), how often you want digests, and how many papers each digest should hold. "Download Models & Start" pulls the weights from HuggingFace and starts the app.

Expect the first download to take a few minutes. Later runs load the cached weights.

Daily use

Sift lives in the macOS menu bar or the Linux system tray.

Menu item What it does
Open Latest Digest Opens the most recent digest in your browser
Digest History Lists past digests inside the retention window
Saved Papers Opens the saved-papers page with your notes
Fetch Now Runs the pipeline immediately
Preferences Topics, schedule, models, retention, notifications
Model Info Shows the models and hardware in use
Pause Until Tomorrow Skips the remaining runs today

When a digest finishes, the tray icon changes and you get a system notification. The digest itself does not steal focus; open it when you are ready.

Inside the digest page, the buttons talk to a small HTTP server bound to 127.0.0.1 on a random free port, authenticated with a token generated on first run and stored in your config. That is what records reads, saves, and notes back into the database.

To change topics, open Preferences, type a topic and click Add, or select one and click Remove. Changes apply on the next fetch. Switching models works the same way: pick from the dropdown and save, and Sift downloads the new weights if they are not cached.

Configuration

Preferences covers everything, but the config file is plain YAML if you prefer editing it directly.

  • macOS: ~/Library/Application Support/sift/config.yaml
  • Linux: ~/.config/sift/config.yaml
Key Default Meaning
topics from wizard List of ArXiv search phrases
digest_frequency once_daily once_daily or twice_daily
digest_hour_morning 8 Hour of the morning run, 24h clock
digest_hour_evening 18 Hour of the evening run when twice daily
max_papers 10 Papers per digest, after dedupe and ranking
hours_back 24 Ignore papers published before this window
summarization_model hardware-based HuggingFace seq2seq model id
embedding_model hardware-based sentence-transformers model name
digest_retention_days 30 Age at which old digests are pruned
notifications_enabled true System notification when a digest is ready
action_port assigned Preferred port for the action server
action_token generated Auth token for the action server

Restart Sift after editing the file by hand.

Model reference

Hardware Summarization model Embedding model
Apple Silicon, under 16 GB sshleifer/distilbart-cnn-6-6 all-MiniLM-L6-v2
Apple Silicon, 16-32 GB sshleifer/distilbart-cnn-12-6 all-MiniLM-L6-v2
Apple Silicon, 32 GB+ facebook/bart-large-cnn all-mpnet-base-v2
Intel Mac sshleifer/distilbart-cnn-6-6 all-MiniLM-L6-v2
Linux, CPU only sshleifer/distilbart-cnn-6-6 all-MiniLM-L6-v2
Linux GPU, under 8 GB VRAM sshleifer/distilbart-cnn-12-6 all-MiniLM-L6-v2
Linux GPU, 8-16 GB VRAM facebook/bart-large-cnn all-mpnet-base-v2
Linux GPU, 16 GB+ VRAM google/pegasus-large all-mpnet-base-v2

These are suggestions, not limits. Any of the four summarization models and two embedding models can be selected in the wizard or in Preferences.

Embeddings are stored with the name of the model that produced them, so switching embedding models re-embeds instead of mixing vectors from different spaces.

How it works

The pipeline runs on a background thread, guarded by a lock so two runs never overlap.

  1. Fetch. Query the ArXiv Atom API once per topic, tagging each result with the topic that found it, and drop anything outside the hours_back window.
  2. Select. Discard papers that already appeared in a digest, then take up to max_papers round-robin across topics.
  3. Embed. Encode abstracts with sentence-transformers and store the vectors in SQLite along with the model name.
  4. Cluster. Run k-means for k from 2 to 5, score each fit by cosine silhouette, keep the best k, and label each cluster with its top TF-IDF terms.
  5. Summarize. Generate a short summary per paper with the seq2seq model on CUDA, MPS, or CPU. If generation fails, fall back to the first two sentences of the abstract.
  6. Render. Write a self-contained HTML page with Jinja2 to ~/Documents/Sift/, record the digest, and prune digests past the retention window.

ArXiv API limits

Sift uses the ArXiv API, which rate limits automated requests.

  • Between topics: at least 5 seconds between requests, in line with ArXiv's terms of service.
  • On HTTP 429: wait 30 seconds and retry, then 60 seconds and retry once more. If it is still blocked, that topic is skipped for this run and an error is logged.
  • Persistent blocks: ArXiv sometimes blocks an IP for hours. Repeated 429s across several runs mean the block is at the IP level and nothing in Sift can clear it. Running many fetches in a short window is the usual cause.

Where your data lives

What macOS Linux
Config ~/Library/Application Support/sift/config.yaml ~/.config/sift/config.yaml
Database ~/Library/Application Support/sift/sift.db ~/.local/share/sift/sift.db
Logs ~/Library/Logs/sift/sift.log ~/.local/state/sift/log/sift.log
Digest pages ~/Documents/Sift/ ~/Documents/Sift/

Digests are written under ~/Documents rather than a data directory so sandboxed browsers, such as Firefox installed as a snap, can open them.

The database is SQLite in WAL mode, with tables for papers, digests, and saved papers. To reset Sift completely, quit it and delete the config file and the database.

Start at login

macOS. Build or install Sift.app in /Applications, then load the bundled LaunchAgent:

cp com.sift.app.plist ~/Library/LaunchAgents/
launchctl load ~/Library/LaunchAgents/com.sift.app.plist

Linux. Copy the desktop entry into your autostart directory:

cp sift.desktop ~/.config/autostart/

The Exec line points at /usr/local/bin/sift. Running from source, change it to your interpreter and main.py.

Development

Install the test dependencies, which deliberately exclude torch and transformers so the suite stays fast:

pip install -r requirements-test.txt
pytest tests/ -v

The suite is 104 tests covering the fetcher, database, clusterer, embedder, summarizer, renderer, scheduler, action server, hardware detection, and pipeline. GitHub Actions runs it on Python 3.11 and 3.12 for every push and pull request to main.

Packaged builds are produced with PyInstaller:

pyinstaller sift.spec

Module layout:

Module Responsibility
app/pipeline.py Orchestrates the five pipeline steps on a worker thread
app/fetcher.py ArXiv client, retries, rate limiting
app/embedder.py sentence-transformers wrapper, lazy model load
app/clusterer.py k-means, silhouette-based k, TF-IDF labels
app/summarizer.py seq2seq summarization with device selection
app/renderer.py Jinja2 rendering of digest and saved pages
app/db.py SQLite schema, migrations, queries
app/action_server.py Token-authenticated localhost server for page actions
app/scheduler.py APScheduler cron jobs and pause handling
app/menubar.py, app/tray_linux.py macOS and Linux tray UIs
app/wizard.py, app/preferences.py Tkinter setup and settings windows
app/hardware.py Chip, memory, and CUDA detection with model recommendations
app/notifier.py System notifications via osascript or notify-send

Troubleshooting

No tray icon on Linux. The default backend is GTK because the older xorg backend needs a legacy system tray that GNOME and Wayland no longer provide. Install the GTK bindings listed under Requirements, or set PYSTRAY_BACKEND to pick another backend. Without a tray, Sift falls back to running the pipeline once from the command line.

"No new papers to digest." Everything ArXiv returned was already in an earlier digest, or nothing matched inside the hours_back window. Widen hours_back or add topics.

Buttons in the digest do nothing. The action server only runs while Sift is running, and old digests carry the port and token they were rendered with. Reopen the latest digest after restarting.

Slow summaries. Summarization dominates the run. Drop to a smaller model in Preferences, or lower max_papers.

Logs go to the path in the table above and to stdout when you run from source.

Contributing

Issues and pull requests are welcome. Areas that would help most:

  • Better cluster labeling than TF-IDF terms
  • Windows support
  • More model profiles
  • Non-ArXiv sources

License

Apache-2.0. See LICENSE.

About

ArXiv digest tool that runs completely locally with summarization models.

Topics

Resources

Stars

6 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages