Create a new virtual environment
conda env create -f environment.ymlpython train.py --prefix data/kmer_ --model ./model/checkpoint.pt --epoch 1000 --batch_size 128 --patience 50 --convert False --fasta ./data/fasta_examples/DNA_sequences.txtprefix: The prefix of the data file used for model training, validation, and testing. The path of example training file is 'data/kmer_pos_train.csv'.
model: The path where the model is to be stored.
epoch: The number of epochs for model training.
batch_size: The batch size for model training.
patience: The early stopping threshold for model training. If the model continues for 'patience' epochs and the loss does not decrease, then the training will stop.
convert: Whether to convert fasta sequences to kmer features. Fasta sequences can only be input into the model for training after being converted into kmer features. If set to True, the convert step is conducted.
fasta: The path of the fasta file to be converted to kmer features. If convert is True, the fasta file should not be null.
python predict.py --threshold 0.5 --path ./data/fasta_examples/DNA_sequences.txt --model ./model/checkpoint.ptthreshold: The threshold for predicting a positive sample. Default is 0.5. If the probability of a sample being predicted as a positive sample is greater than the threshold, then this sample is considered a positive sample.
path: The path of the DNA fasta file. The format of the file content must strictly adhere to the fasta format, starting a fasta sample sequence with a ">" and a sequence name.
model: The path where the trained model is located.