perplexity_evals
perplexity_evals
¶
Perplexity evaluations for existing datasets with validation splits.
These complement the existing RolloutEvaluation counterparts by measuring how well the model predicts validation-set tokens (1/perplexity, higher is better), which is especially useful for tracking forgetting in continual learning experiments.