home / grpo-gsm8k-test / summaries

summaries: 69

This data as json

id task_name model_tag total_examples correct accuracy no_answer_count stop_reason_counts duration_human pass_k temperature top_p max_tokens error model
69 TOTAL 34 models 214846 125885 0.585931318246558 21764 {} 1h 45m 38s 1 0.0 0.95 1024   Models: grpo-gsm8k-test-30ep-success_global_step_50, grpo-gsm8k-test-30ep-success_global_step_100, grpo-gsm8k-test-30ep-success_global_step_150, grpo-gsm8k-test-30ep-success_global_step_200, grpo-gsm8k-test-30ep-success_global_step_250, grpo-gsm8k-test-30ep-success_global_step_300, grpo-gsm8k-test-30ep-success_global_step_350, grpo-gsm8k-test-30ep-success_global_step_400, grpo-gsm8k-test-30ep-success_global_step_450, grpo-gsm8k-test-30ep-success_global_step_500, grpo-gsm8k-test-30ep-success_global_step_550, grpo-gsm8k-test-30ep-success_global_step_600, grpo-gsm8k-test-30ep-success_global_step_650, grpo-gsm8k-test-30ep-success_global_step_700, grpo-gsm8k-test-30ep-success_global_step_750, grpo-gsm8k-test-30ep-success_global_step_800, grpo-gsm8k-test-30ep-success_global_step_850, grpo-gsm8k-test-30ep-success_global_step_900, grpo-gsm8k-test-30ep-success_global_step_950, grpo-gsm8k-test-30ep-success_global_step_1000, grpo-gsm8k-test-30ep-success_global_step_1050, grpo-gsm8k-test-30ep-success_global_step_1100, grpo-gsm8k-test-30ep-success_global_step_1150, grpo-gsm8k-test-30ep-success_global_step_1200, grpo-gsm8k-test-30ep-success_global_step_1250, grpo-gsm8k-test-30ep-success_global_step_1300, grpo-gsm8k-test-30ep-success_global_step_1350, grpo-gsm8k-test-30ep-success_global_step_1400, grpo-gsm8k-test-30ep-success_global_step_1450, grpo-gsm8k-test-30ep-success_global_step_1500, grpo-gsm8k-test-30ep-success_global_step_1550, grpo-gsm8k-test-30ep-success_global_step_1600, grpo-gsm8k-test-30ep-success_global_step_1620, Qwen_Qwen2.5-1.5B-Instruct | Tasks: gsm8k_main(0), hendrycks_math(0)
Powered by Datasette · Queries took 0.622ms