A local ArXiv research digest for macOS and Linux. Sift fetches new papers on the topics you follow, summarizes them with a HuggingFace model running on your machine, groups them by theme, and hands you a single HTML page to read.
Built in the hopes of saving time for researchers wanting to stay up-to-date with preprints.
- Features
- Requirements
- Installation
- First run
- Daily use
- Configuration
- Model reference
- How it works
- ArXiv API limits
- Where your data lives
- Start at login
- Development
- Troubleshooting
- Contributing
- License
- Local summarization. Abstracts are condensed by a transformer model on your own hardware. Offline after the first model download.
- Topic clustering. Papers are embedded and grouped by k-means, with the number of clusters chosen by a cosine silhouette sweep instead of a fixed k. Each cluster gets a TF-IDF label.
- Fair topic mixing. Papers are picked round-robin across your topics, so one busy topic cannot crowd out the rest.
- No repeats. Papers that already appeared in an earlier digest are skipped.
- Read, save, and annotate. Every paper in the digest has Mark Read, Save, and Copy BibTeX buttons. Saved papers get their own page with a notes field.
- Digest history. Past digests stay in the menu until they age out of the retention window.
- Scheduled or on demand. Once or twice a day at hours you choose, plus a Fetch Now item and a pause-until-tomorrow switch.
- Hardware-aware defaults. The setup wizard detects your chip, memory, and CUDA devices and picks model sizes that fit.
- Python 3.11 or later
- macOS 12+ or Linux x86_64
- 4-16 GB free disk space, depending on the model you pick
- Linux tray support needs GTK bindings:
sudo apt install python3-gi gir1.2-ayatanaappindicator3-0.1
git clone https://github.com/sukhleenk/sift.git
cd sift
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtpython main.pyRunning Sift from source is currently more stable than the packaged builds in
dist/, and is the recommended path.
A setup wizard opens on first launch. It reports the hardware it found, suggests a summarization and embedding model, and lets you override both. You also set your topics (comma separated), how often you want digests, and how many papers each digest should hold. "Download Models & Start" pulls the weights from HuggingFace and starts the app.
Expect the first download to take a few minutes. Later runs load the cached weights.
Sift lives in the macOS menu bar or the Linux system tray.
| Menu item | What it does |
|---|---|
| Open Latest Digest | Opens the most recent digest in your browser |
| Digest History | Lists past digests inside the retention window |
| Saved Papers | Opens the saved-papers page with your notes |
| Fetch Now | Runs the pipeline immediately |
| Preferences | Topics, schedule, models, retention, notifications |
| Model Info | Shows the models and hardware in use |
| Pause Until Tomorrow | Skips the remaining runs today |
When a digest finishes, the tray icon changes and you get a system notification. The digest itself does not steal focus; open it when you are ready.
Inside the digest page, the buttons talk to a small HTTP server bound to 127.0.0.1 on a random free port, authenticated with a token generated on first run and stored in your config. That is what records reads, saves, and notes back into the database.
To change topics, open Preferences, type a topic and click Add, or select one and click Remove. Changes apply on the next fetch. Switching models works the same way: pick from the dropdown and save, and Sift downloads the new weights if they are not cached.
Preferences covers everything, but the config file is plain YAML if you prefer editing it directly.
- macOS:
~/Library/Application Support/sift/config.yaml - Linux:
~/.config/sift/config.yaml
| Key | Default | Meaning |
|---|---|---|
topics |
from wizard | List of ArXiv search phrases |
digest_frequency |
once_daily |
once_daily or twice_daily |
digest_hour_morning |
8 |
Hour of the morning run, 24h clock |
digest_hour_evening |
18 |
Hour of the evening run when twice daily |
max_papers |
10 |
Papers per digest, after dedupe and ranking |
hours_back |
24 |
Ignore papers published before this window |
summarization_model |
hardware-based | HuggingFace seq2seq model id |
embedding_model |
hardware-based | sentence-transformers model name |
digest_retention_days |
30 |
Age at which old digests are pruned |
notifications_enabled |
true |
System notification when a digest is ready |
action_port |
assigned | Preferred port for the action server |
action_token |
generated | Auth token for the action server |
Restart Sift after editing the file by hand.
| Hardware | Summarization model | Embedding model |
|---|---|---|
| Apple Silicon, under 16 GB | sshleifer/distilbart-cnn-6-6 | all-MiniLM-L6-v2 |
| Apple Silicon, 16-32 GB | sshleifer/distilbart-cnn-12-6 | all-MiniLM-L6-v2 |
| Apple Silicon, 32 GB+ | facebook/bart-large-cnn | all-mpnet-base-v2 |
| Intel Mac | sshleifer/distilbart-cnn-6-6 | all-MiniLM-L6-v2 |
| Linux, CPU only | sshleifer/distilbart-cnn-6-6 | all-MiniLM-L6-v2 |
| Linux GPU, under 8 GB VRAM | sshleifer/distilbart-cnn-12-6 | all-MiniLM-L6-v2 |
| Linux GPU, 8-16 GB VRAM | facebook/bart-large-cnn | all-mpnet-base-v2 |
| Linux GPU, 16 GB+ VRAM | google/pegasus-large | all-mpnet-base-v2 |
These are suggestions, not limits. Any of the four summarization models and two embedding models can be selected in the wizard or in Preferences.
Embeddings are stored with the name of the model that produced them, so switching embedding models re-embeds instead of mixing vectors from different spaces.
The pipeline runs on a background thread, guarded by a lock so two runs never overlap.
- Fetch. Query the ArXiv Atom API once per topic, tagging each result with the topic that found it, and drop anything outside the
hours_backwindow. - Select. Discard papers that already appeared in a digest, then take up to
max_papersround-robin across topics. - Embed. Encode abstracts with sentence-transformers and store the vectors in SQLite along with the model name.
- Cluster. Run k-means for k from 2 to 5, score each fit by cosine silhouette, keep the best k, and label each cluster with its top TF-IDF terms.
- Summarize. Generate a short summary per paper with the seq2seq model on CUDA, MPS, or CPU. If generation fails, fall back to the first two sentences of the abstract.
- Render. Write a self-contained HTML page with Jinja2 to
~/Documents/Sift/, record the digest, and prune digests past the retention window.
Sift uses the ArXiv API, which rate limits automated requests.
- Between topics: at least 5 seconds between requests, in line with ArXiv's terms of service.
- On HTTP 429: wait 30 seconds and retry, then 60 seconds and retry once more. If it is still blocked, that topic is skipped for this run and an error is logged.
- Persistent blocks: ArXiv sometimes blocks an IP for hours. Repeated 429s across several runs mean the block is at the IP level and nothing in Sift can clear it. Running many fetches in a short window is the usual cause.
| What | macOS | Linux |
|---|---|---|
| Config | ~/Library/Application Support/sift/config.yaml |
~/.config/sift/config.yaml |
| Database | ~/Library/Application Support/sift/sift.db |
~/.local/share/sift/sift.db |
| Logs | ~/Library/Logs/sift/sift.log |
~/.local/state/sift/log/sift.log |
| Digest pages | ~/Documents/Sift/ |
~/Documents/Sift/ |
Digests are written under ~/Documents rather than a data directory so sandboxed browsers, such as Firefox installed as a snap, can open them.
The database is SQLite in WAL mode, with tables for papers, digests, and saved papers. To reset Sift completely, quit it and delete the config file and the database.
macOS. Build or install Sift.app in /Applications, then load the bundled LaunchAgent:
cp com.sift.app.plist ~/Library/LaunchAgents/
launchctl load ~/Library/LaunchAgents/com.sift.app.plistLinux. Copy the desktop entry into your autostart directory:
cp sift.desktop ~/.config/autostart/The Exec line points at /usr/local/bin/sift. Running from source, change it to your interpreter and main.py.
Install the test dependencies, which deliberately exclude torch and transformers so the suite stays fast:
pip install -r requirements-test.txt
pytest tests/ -vThe suite is 104 tests covering the fetcher, database, clusterer, embedder, summarizer, renderer, scheduler, action server, hardware detection, and pipeline. GitHub Actions runs it on Python 3.11 and 3.12 for every push and pull request to main.
Packaged builds are produced with PyInstaller:
pyinstaller sift.specModule layout:
| Module | Responsibility |
|---|---|
app/pipeline.py |
Orchestrates the five pipeline steps on a worker thread |
app/fetcher.py |
ArXiv client, retries, rate limiting |
app/embedder.py |
sentence-transformers wrapper, lazy model load |
app/clusterer.py |
k-means, silhouette-based k, TF-IDF labels |
app/summarizer.py |
seq2seq summarization with device selection |
app/renderer.py |
Jinja2 rendering of digest and saved pages |
app/db.py |
SQLite schema, migrations, queries |
app/action_server.py |
Token-authenticated localhost server for page actions |
app/scheduler.py |
APScheduler cron jobs and pause handling |
app/menubar.py, app/tray_linux.py |
macOS and Linux tray UIs |
app/wizard.py, app/preferences.py |
Tkinter setup and settings windows |
app/hardware.py |
Chip, memory, and CUDA detection with model recommendations |
app/notifier.py |
System notifications via osascript or notify-send |
No tray icon on Linux. The default backend is GTK because the older xorg backend needs a legacy system tray that GNOME and Wayland no longer provide. Install the GTK bindings listed under Requirements, or set PYSTRAY_BACKEND to pick another backend. Without a tray, Sift falls back to running the pipeline once from the command line.
"No new papers to digest." Everything ArXiv returned was already in an earlier digest, or nothing matched inside the hours_back window. Widen hours_back or add topics.
Buttons in the digest do nothing. The action server only runs while Sift is running, and old digests carry the port and token they were rendered with. Reopen the latest digest after restarting.
Slow summaries. Summarization dominates the run. Drop to a smaller model in Preferences, or lower max_papers.
Logs go to the path in the table above and to stdout when you run from source.
Issues and pull requests are welcome. Areas that would help most:
- Better cluster labeling than TF-IDF terms
- Windows support
- More model profiles
- Non-ArXiv sources
Apache-2.0. See LICENSE.
