Skip to content

make column names configurable #36

Description

@jo-tham
As a user
I want to be able to specify which columns correspond to which gypsy variables
So that I can run gypsy without renaming the fields in my data

Currently, the whole system assumes column names are fixed to our constants, e.g. here

This is inconvenient for users as described in user story above.

We can make it configurable by indexing the data using global variables, e.g.

data[COLUMN_NAME]

Configuring a bunch of module level constants (singletons) is difficult and dangerous. this is a good place to use a class with getters and setters as here for example

This will store all the columns as properties on a class so we can do

data[col_config.black_spruce_age]

And those properties will be protected from modification by overriding the setter as shown in the link above

the drawback here is that this col_config class must be used anywhere we index the data. it may not be intuitive for new developers


  • check for existing implementations - config file library may do the trick here
  • setup sensible default config
  • refactor
    • determine optimal way to pass the config around (e.g. as a function param?)
    • maybe pandas already has a smart way to do this, it would be nice to pass it around with the data because passing it around as a function param - it will have to go to many functions in the library
    • compare which library functions actually need the config
      • which use a data frame and which simply use values?
    • refactor all column references to use the config system

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions