This README explains how to use the YAML configuration file to set up dataset loading, model architecture, training, and prediction parameters for the Transformer regression model.
Example config files can be found in the tests folder.
-
input_files
Expects an array of.h5files. -
scaling_dict
A dictionary specifying how each variable is scaled. Available scalers include:
log,logminmax,minmax,standard,arctan,tanh, and custom physical scalers.
More details on scalers can be found insrc/data/DataScalers.py.
All data before and after scaling is plotted.
Use"none"to apply no scaling. Variables not listed will be excluded from the model. -
target_scaler
Specifies how to scale the target variable; plots are also generated. -
trainandvalbatch_size: Number of events per batchnum_workers: Number of CPU workers used for data loading
-
inverse_sampling: IfTrue, samples batches unevenly to prioritize rare events by building a KDE with specified parameters and using its inverse distribution. -
KDE_width: Width parameter for Gaussian KDE. -
Min_Dens_cap: Caps minimum density to prevent excessively large sampling weights for ultra-rare events.
-
type
Specify model type:DNNRegressionorTransformerRegression. -
name
String identifying the model. This becomes the name of the results directory containing all generated plots. -
n_particles
Number of particles in the data, ordered as (lep, tau, MET, jet1, jet2, jet3, etc.).
The model adjusts automatically if more than 5 particles are used and applies appropriate masking.
layers
Array specifying hidden layer sizes.
Example:[128, 128]means two hidden layers with 128 neurons each.
-
num_heads
Number of attention heads. -
num_layers
Number of stacked transformer layers. -
input_embedder_NN_layers
Array defining neuron counts in the input embedder layers. The last value is the embedding dimension and must be divisible by the number of heads. -
output_activation
Activation function for output neuron:softplus,sigmoid,tanh,relu, orlinear(no activation). -
pooling_type
Eithermeanorattention. -
norm_type
Normalization method between input embedder layers:none,batchnorm, orlayernorm.
-
dropout_probability
Dropout rate applied to hidden layers. -
early_stopping_patience
Number of epochs with no improvement before stopping training. -
learning_rate
Learning rate for optimizer. -
loss_fn
Loss function name. -
loss_fn_params
Dictionary of parameters to configure the loss function.
Supports nested loss functions, e.g.:
loss_fn: QuantileAwareLoss
loss_fn_params:
alpha: 0.1
squared: false
base_loss: SmoothL1Loss
base_loss_params:
beta: 10.0
reduction: meanAvailable loss functions are listed in src/utils/LossFunctions.py.
-
lr_scheduler_patience
Number of epochs to wait before reducing learning rate when validation loss plateaus (default scheduler:PlateauReduceLR). -
n_epochs
Total number of training epochs. -
num_quantiles
Number of quantiles used when employingQuantileAwareLoss. Ignored otherwise. -
weight_decay
L2 regularization strength. Set to0to disable. -
optimize
Enables PyTorch TF32 matmul optimization (trades precision for speed). -
Transformer-specific:
compute_interaction_tokens: Set toTrueto add interaction tokens between particles as described in [10.1088/1674-1137/ad7f3d]. Adds a separate interaction embedder with layers matching the particle embedder.
-
device
e.g.,cudaorcpu. -
mode
Model mode:train,performance, orpredict. -
runner
Identifier for the runner (e.g.,Bob).
-
Statistical such as: MeanAbsoluteError, MeanSquaredError, R2Score, MeanFractionalBias, MedianAbsoluteError and RootMeanSquaredLogError are calculated for the DNN and the Transformer.
-
Extra attention-specific metrics are computed for the Transformer: AttentionEntropy, AttentionSparsity, HeadDiversity
-
Feature importance plots and performance-related plots are generated and saved in the model directory.
-
When KDE sampling is enabled, plots comparing original vs. new sampling distributions are included.