Policy on finetuning on benchmark train splits

#140
by sz14 - opened

Hello, I wanted to ask if there is any policy against training on the "train" splits of any benchmark listed here (NOT validation or test sets).
MobileLLM R1 140M already does this from what I can glean from their card. So do Nexus-Erebus and Isabel models.

Asking because I did try it and saw a big boost across the board mixing them with nemotron-cc, even on benchmarks whose training splits I did not include. I need a strong backbone for my uses anyway so I will be keeping that checkpoint, but want to confirm before I update anything here.

sz14 changed discussion title from Policy on training on benchmark train splits to Policy on finetuning on benchmark train splits
Axiomic Labs org

Hi!
In terms of Fine tuning on benchmark splits, we consider that benchmaxxing. And in terms of pretraining, we do advise against using them as it makes the benchmarks no longer held out evals.
While we do make exceptions for models that aren't toward the top/use it for a tiny part of their pretraining corpus, it is generally not advised (nor is it necessary, for example none of the models in the top 5 of any category use train splits).

Any other questions lmk!

No that is all, thank you! Well ig this would still be pre-training phase, since Instruct-SFT was planned after that, but regardless it does not matter.

I will not be updating the model data here then, but will update my repo. I am making a VLM and this nemotron (longer context) + training mix pre-train extension is a much better base than the one listed here irrespective of scores (I also test on actual generation on a few handwritten examples over a variety of topics).

Just letting you know to account for the drift!

for example none of the models in the top 5 of any category use train splits

This is technically true (falls apart at top 6), but on that note I want to clarify this for Instruct/Chat models as well.
smol-smoltalk and openhermes 2.5 are commonly used datasets and from what I can tell both of them have the train sets of these benchmarks in them, included from their constituent datasets.

So would I be able to release the Instruct model or would even that be flagged? Pretty much following the SmolLM2-Instruct recipe.

sz14 changed discussion status to closed
Axiomic Labs org

Hmm, from looking into it the open Hermes section of smol smoltalk contains responses generated from the train splits of GSM8K train split and MATH train split. However, since we dont evaluate on these, shouldnt be an issue!

I saw a model on HN that did this: https://mvakde.github.io/blog/44-on-arc-1/

I saw a model on HN that did this: https://mvakde.github.io/blog/44-on-arc-1/

Already doing this with the next release, difference though is that it might be good at tasks like that + being good at language modeling. At least that's our hope.

Sign up or log in to comment