Skip to content

Latest commit

 

History

22 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Instructions for Accessing Records Relating to Membership in the Nationalsozialistische Deutsche Arbeiterpartei (NSDAP) (A3340) on the AWS Registry of Open Data

The National Archives and Records Administration (NARA), in partnership with Amazon Web Services (AWS) Open Data Sponsorship, posted two portions (MFKL and MFOK) of Records Relating to Membership in the Nationalsozialistische Deutsche Arbeiterpartei (National Socialist German Labor Party (NSDAP)), 1927-1945 (A3340) (National Archives Identifier 12044361) to the AWS Registry of Open Data. The membership cards in the Zentralkartei (MFKL) records comprise the alphabetical membership registry maintained by the NSDAP in its central administrative offices. The membership cards in the Ortsgruppenkartei (MFOK) records provide a central geographical registry of Nazi Party members.

This documentation guides users in how to access the data.

About the Dataset

The NSDAP dataset on the AWS Registry of Open Data includes over 14 million digital objects and metadata, including Textract-generated OCR text. The files are arranged into directories by microfilm roll, with each directory containing a TIF file for each image on the roll, a PDF that combines all of the images on the roll, and a JSON file with the metadata about the roll and extracted text from each image. The OCR was generated by the National Archives Catalog using Amazon Textract.

The fields in the JSON file reflect the structure of the Catalog records that contain these images; i.e. "title" is the title of the file unit, "naId" is the National Archives Identifier, "containerId" is the microfilm roll number. The OCR is stored in the "extractedText" field for each digital object, and the "objectFilename" maps the OCR information to the specific image within the directory.

Access Methods

The AWS Registry of Open Data is a service provided by AWS to store open, public datasets for free so that they can be accessed and analyzed on AWS. Users can access both the full dataset and specific portions of the dataset using the AWS Command Line Interface (CLI) , an open source tool that enables users to interact with AWS services using commands in their command-line. Documentation for AWS CLI is available here.

Accessing the Full Dataset

The full dataset can be accessed with the following Amazon Resource Name (ARN): arn:aws:s3:::nara-nsdap

The following AWS CLI command will list the full dataset:

aws s3 ls s3://nara-nsdap/ --no-sign-request

The following AWS CLI command will download the full dataset:

aws s3 sync s3://nara-nsdap/ [destination] --no-sign-request

Note

Please be advised the full dataset is very large and may require external storage.

Accessing Portions of the Dataset

If you would like to download portions of the dataset, you may use the AWS CLI command builder. This tool is sorted into a searchable table first by the Zentralkartei (MFKL) records and Ortsgruppenkartei (MFOK) records and then alphabetically by last name. You may sort on any of the columns in the table including National Archives ID, Collection, Box, and Description (Name). You may also expand the table to view up to 25 entries per page. Clicking ‘Use Path’ on any row will populate the ‘Generated Command’ field at the bottom of the page. You may click to check specific file types you would like to exclude from your download including JSON, TIF, and PDF. Click ‘Copy Command’ to copy the AWS CLI terminal command to your dashboard, and then paste it into your command shell to initiate the download.

Potential Use Cases

  1. Download a single microfilm roll Instead of syncing the entire 14+ million object dataset, users can download the contents of a specific microfilm roll, which includes its TIF images, the combined PDF, and the JSON metadata file.

  2. Isolate and download the JSON metadata To perform text analysis or search the OCR across the entire collection without downloading terabytes of images, users can filter their sync to download only the JSON files.

  3. Search OCR text across rolls Once the JSON files are downloaded locally, users can search the extractedText fields for specific names, towns, or occupations.

Note

Note: For the Description field in the Bulk Download AWS S3 CLI Command Builder for NSDAP, text extraction from the records was performed with the assistance of AI and reviewed by NARA staff.

image

About

Instructions for Accessing Records Relating to Membership in the Nationalsozialistische Deutsche Arbeiterpartei (NSDAP) (A3340) on the AWS Registry of Open Data

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors