Hi
I am the maintainer of another spacy pipeline sentiment library and i am trying to figure out how to benchmark spacy sentiment models fairly.
i have written something here https://github.com/sloev/sentimental-onix/tree/main/benchmark
it uses this dataset https://archive.ics.uci.edu/ml/datasets/Sentiment+Labelled+Sentences as foundation for a benchmark.
my issue is that both spacytextblob and my library outputs floating points but in order to validate against a test dataset i am trying to threshold our values into descrete labels neg, neu, pos.
but whether it turns out to be a fair comparison is hard for me to evaluate.
results as they are (my model uses Onnx based sentiment model, and a default threshold of neg < -0.7 < neu < 0.7 < pos)
are:
| library |
result |
| spacytextblob |
58.9% |
| sentimental_onix |
69% |
kind regards
Hi
I am the maintainer of another spacy pipeline sentiment library and i am trying to figure out how to benchmark spacy sentiment models fairly.
i have written something here https://github.com/sloev/sentimental-onix/tree/main/benchmark
it uses this dataset https://archive.ics.uci.edu/ml/datasets/Sentiment+Labelled+Sentences as foundation for a benchmark.
my issue is that both spacytextblob and my library outputs floating points but in order to validate against a test dataset i am trying to threshold our values into descrete labels neg, neu, pos.
but whether it turns out to be a fair comparison is hard for me to evaluate.
results as they are (my model uses Onnx based sentiment model, and a default threshold of neg < -0.7 < neu < 0.7 < pos)
are:
kind regards