Skip to content
DCC-BSPublic

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

1 watching

Forks

Repository files navigation

README.md


Setup

Add a .env file in the root folder with the following content:

Email Configuration

DATASPOT_EMAIL_RECEIVERS_TECHNICAL_ONLY=["petra.muster@bs.ch", "peter.muster@bs.ch"]
DATASPOT_EMAIL_RECEIVERS=["petra.muster@bs.ch", "peter.muster@bs.ch"]
DATASPOT_EMAIL_SERVER=
DATASPOT_EMAIL_SENDER=

Authentication Configuration

M2M Authentication (Recommended)

For machine-to-machine authentication using Azure AD/Entra ID:

# Azure AD/Entra ID configuration
DATASPOT_TENANT_ID=your-azure-tenant-id
DATASPOT_CLIENT_ID=your-entra-app-client-id
DATASPOT_CLIENT_SECRET=your-entra-app-client-secret
DATASPOT_EXPOSED_CLIENT_ID=5d25a...612

# Dataspot service user access key
DATASPOT_SERVICE_USER_ACCESS_KEY=your-service-user-access-key

Legacy Username/Password Authentication (Deprecated for DataspotAuth)

The following environment variables can be deleted, if they still exist. They are from a legacy system.

DATASPOT_EDITOR_USERNAME
DATASPOT_EDITOR_PASSWORD
DATASPOT_ADMIN_USERNAME
DATASPOT_ADMIN_PASSWORD

DATASPOT_CLIENT_ID
DATASPOT_AUTHENTICATION_TOKEN_URL
DATASPOT_API_BASE_URL

Note: The authentication system has been updated to use M2M authentication. The legacy username/password authentication may be deprecated in the future.


OGD Dataset Sync: Restricted vs Public

OGD datasets from Huwise (ODS) are synced into Dataspot by two independent scripts, run in this order:

  1. scripts/sync_ods_restricted_datasets.py - syncs datasets with is_restricted=True as DRAFT (WORKING) into OGD-Datensätze aus Huwise (unveröffentlicht), with no compositions, Huwise deployment, or OGD distributions.
  2. scripts/sync_ods_datasets.py (and scripts/sync_ods_dataset_compositions.py) - syncs datasets with is_restricted=False as PUBLISHED into OGD-Datensätze aus Huwise, and also promotes any dataset that has left restriction.

There are three sibling collections under DCC Data Competence Center:

  • OGD-Datensätze aus Huwise - published datasets.
  • OGD-Datensätze aus Huwise (unveröffentlicht) - newly synced restricted datasets land here.
  • OGD-Datensätze aus Huwise (intern) - stewards manually move datasets here that must never be auto-published (e.g. internal/test datasets).
flowchart TD
    odsListing["ODS Automation API listing\n(dataset_id + is_restricted)"]
    odsListing --> isRestricted{"is_restricted?"}
    isRestricted -->|false| publicScript["sync_ods_datasets.py\n(own is_restricted=False fetch)"]
    isRestricted -->|true| restrictedScript["sync_ods_restricted_datasets.py\n(full listing)"]

    restrictedScript --> existsCheck{"exists in Dataspot?"}
    existsCheck -->|no| createDraft["Create Dataset\nstatus=WORKING\nunveroeffentlicht folder"]
    existsCheck -->|yes| statusCheck{"current status?"}
    statusCheck -->|PUBLISHED or DELETENEW| deferMain["Skip - owned by\npublicScript / human"]
    statusCheck -->|WORKING| compareUpdate["sync_datasets\nstatus=WORKING\nno deployments/distributions\n(folder preserved)"]

    publicScript --> mappingCheck{"already exists\nin Dataspot?"}
    mappingCheck -->|no| createPublished["Create Dataset\nstatus=PUBLISHED\nmain folder"]
    mappingCheck -->|yes| statusGate{"current status\nWORKING?"}
    statusGate -->|"no (already PUBLISHED, etc.)"| normalUpdate["Normal update\n(existing behavior)"]
    statusGate -->|yes| folderCheck{"current folder\nUUID?"}
    folderCheck -->|intern| skipInternal["Exclude from sync.\nLog + email as skipped-internal"]
    folderCheck -->|unveröffentlicht| promoteMove["Script pre-pass:\nmove to main folder\nset status=PUBLISHED"]
    folderCheck -->|anywhere else| promoteInPlace["Script pre-pass:\nset status=PUBLISHED\nleave folder unchanged"]
    promoteMove --> normalUpdate
    promoteInPlace --> normalUpdate
Loading

Datasets that are genuinely removed from ODS (neither restricted nor public) are handled by status ownership: sync_ods_restricted_datasets.py permanently deletes WORKING datasets, while sync_ods_datasets.py marks PUBLISHED datasets as DELETENEW. Demotion (a published dataset becoming restricted again) is not handled automatically.


Managing (Data Owner) Posts

The following gif shows everything needed to create a new post, and link the correct person. Note that we don't actually need to create the person or user, as this happens daily automatically. Setting up a new data owner


Integrating code from a dev (or feature) environment into prod [Work-In-Progress]

When integrating a dev into prod, first we need to clone the dev into an int database.

Then:

  1. Export DNK from dev as xlsx and import it again (dry run is enough).
  2. If we don't fix warnings or errors that occur, then they will appear later again.
  3. Integrate yaml from dev into int
  4. Run job "Regelverletzungen prüfen"
  5. Export DNK as xlsx and import it again (dry run is enough)
  6. Export and reimport other models that might be affected aswell
  7. Merge dev into main and delete dev branch

If everything worked without errors, we can apply the int yaml into the prod yaml and reapply the changes made to the int to the prod.

After that, delete the dev branch on github, in dataspot, and also its corresponding Annotations.yaml. Also delete the int environment in dataspot.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages