<-back to projects
COMPLETED2023-09 → 2023-12

CompSem -- NLI

How do different sentence encoders trade off NLI accuracy and transfer performance?

NLPPythonPyTorchDeep Learning

I wanted a clean comparison of sentence encoders beyond single benchmark scores. I built a modular pipeline to train multiple encoders on SNLI and evaluate transfer with SentEval, including reusable checkpoints. The key finding was that higher NLI accuracy did not guarantee better transfer. I handled modeling, training, evaluation, and the write-up. The repository is here: Repo.