(🛠️This is incomplete code currently undergoing testing.)
This code is designed to collect papers from Scopus and is intended for use in a simple project. The Scopus API keys personally issued to you are required.
The pipeline follows a two-stage collection strategy.
-
The Search stage performs broad record collection. It loads query keywords from
keywords.txt, combines them with year-based query chunks, and submits them to Scopus usingScopusSearch. -
This step is designed to maximize collection efficiency. Instead of requesting detailed metadata for every paper immediately, it first saves only the core search results. Each unique paper is stored in
scopus_records.csv, while keyword-to-paper relationships are stored inscopus_matches.csv. -
This design keeps the first-pass collection lightweight and reduces unnecessary API usage.
-
The Enrichment stage performs record-level metadata expansion. It reads the previously collected EIDs from
scopus_records.csvand usesAbstractRetrievalto fetch additional fields. -
This step is intended for metadata that is more expensive to collect, such as abstracts, references, affiliations, and funding information. Because it runs after the Search stage, the user can decide whether detailed enrichment is necessary before spending additional API quota.
-
The Pipeline wrapper controls the overall execution flow. It provides a single entry point for the full workflow while keeping the Search stage and Enrichment stage logically separated.
-
This makes the code easier to maintain, easier to debug, and easier to reuse for different collection settings.
cf. Why separate Search and Enrichment?
: Scopus search results can be collected relatively efficiently, but detailed record retrieval is more expensive and slower. Separating the workflow into two stages makes large-scale collection more efficient, reduces unnecessary API usage, and allows the user to enrich only the records that are actually needed.