Skip to content

Train command

Johann A. Briffa edited this page Nov 23, 2021 · 1 revision

Description

Train a new segmentation model.

Overview diagram

Train command overview diagram

Command arguments

Argument Type Description Required?
--preproc_volume_fullfname full file name Full path to preprocessed volume data file (*.hdf). yes
--subvolume_dir directory Directory of subvolume that was labelled. yes
--label_dirs list of directories Directories of labels containing the slices in subvolume_dir with unlabelled regions blacked out. yes
--config_fullfname full file name Full path to configuration file specifying how to extract features and classify them (*.json). yes
--result_segmenter_fullfname full file name Full path to segmenter pickle file to be created by this process (*.pkl). yes
--trainingset_file_fullfname full file name Full path to file that is used to store the training set (*.hdf). Note that if this is left out then there is nothing to checkpoint. no
--verbose_training yes/no Whether to show sklearn's verbose messages during training (default is yes). no
--train_sample_seed whole number Seed for the random number generator which samples voxels. If left out then the random number generator will be non-deterministic. no
--checkpoint_fullfname full file name Full path to file that is used to let the process save its progress and continue from where it left off in case of interruption (*.json). If left out then the process will run from beginning to end without saving any checkpoints. no
--checkpoint_namespace string Unique name for the group of checkpoints used by this command. no
--reset_checkpoint yes/no Whether to clear the checkpoints about this command from the checkpoint file and start afresh or not (default is no). no
--log_file_fullfname full file name Full path to file that is used to store a log of what is displayed on screen (*.txt). no
--max_processes_featuriser whole number Maximum number of parallel processes to use concurrently whilst featurising (-1 to use maximum, default). no
--max_processes_classifier whole number Maximum number of parallel processes to use concurrently whilst classifying (-1 to use maximum, default). no
--max_batch_memory fractional number Maximum amount of GB to allow for processing the volume in batches (-1 to use maximum, default). no
--use_gpu yes/no Whether to use the GPU for computing features (default is no). no
--print_output yes/no Whether to output to the screen (default is yes). no
--debug_mode yes/no Whether to give full error messages (default is no). no

Command line example

python ASEMI-segmenter/Python/asemi_segmenter/bin/train.py \
    --preproc_volume_fullfname "output/preprocess/volume.hdf" \
    --subvolume_dir "training_set/subvolume" \
    --label_dirs \
        "training_set/labels/air" \
        "training_set/labels/tissues" \
        "training_set/labels/bones" \
    --config_fullfname "output/tune/best_result.json" \
    --result_segmenter_fullfname "output/train/segmenter.pkl" \
    --trainingset_file_fullfname "output/train/trainingset.hdf" \
    --verbose_training "yes" \
    --train_sample_seed 0 \
    --checkpoint_fullfname "output/checkpoint.json" \
    --log_file_fullfname "output/log.txt" \
    --max_processes_featuriser 4 \
    --max_processes_classifier 4 \
    --max_batch_memory 1.0 \
    --use_gpu "no"

Stages

# Log message Description
1 Loading data Load input data, validate it, and initialise the checkpoint and hash function.
2 Hashing subvolume slices Create hash vectors for each slice in the subvolume.
3 Constructing labels dataset Load the labels of all the slices in the training set into memory.
4 Constructing training set Create an intermediate training set HDF file or arrays in memory.
5 Training segmenter Train the segmentation model.

Checkpoints

# Checkpoint name In stage Description
1 creating_trainingset 4 Training set file has been created and should not be overwritten with a new one.
2 constructing_labels 4 Labels have been saved into the training set file.
3 constructing_features_prog 4 Features of all slices in the training set up to the value given have been saved into the training set file.
4 constructing_training_set 4 The training set construction has been completed.
5 training 5 Training has been completed and trained segmenter saved.
6 overall - The process has been completed and does not need to be repeated.

Further information

Clone this wiki locally