Skip to case study
Adhik AdhikariSelected work

Project 04

Fine-Tuned Soccer Discourse Classifier

TakeMeter

TakeMeter tests whether a small fine-tuned language model can distinguish three kinds of discourse inside the same soccer community.

TakeMeter confusion matrix showing correct predictions and errors across analysis, hot take, and reaction labels
Fine-tuned DistilBERT confusion matrix from the project repository.
220
reviewed comments
75.8%
DistilBERT accuracy
81.8%
zero-shot baseline

Dataset and model

I manually reviewed 220 r/soccer comments labeled as analysis, hot take, or reaction, then created stratified training, validation, and test splits. DistilBERT was fine-tuned on a GPU and compared with a zero-shot Llama baseline on the same test set.

Evaluation

The evaluation measured accuracy, precision, recall, F1, and confusion-matrix errors. DistilBERT reached 75.8% accuracy while the zero-shot baseline reached 81.8%, so fine-tuning did not outperform the baseline.

What the errors revealed

The model separated analysis and reaction more reliably than hot takes. Hot-take examples often borrow tactical language from analysis or emotional language from reactions, while being defined mainly by the absence of meaningful evidence. The confusion matrix makes that weak boundary visible.

Next iteration

The documented next step is not a larger claim—it is better data. More hard-negative hot-take examples would give the model clearer evidence for the class boundary that the first dataset did not capture well enough.

Working stack

  • Python
  • Hugging Face
  • DistilBERT
  • Fine-Tuning
  • Model Evaluation