summaries: 83
This data as json
| id | task_name | model_tag | total_examples | correct | accuracy | no_answer_count | stop_reason_counts | duration_human | pass_k | temperature | top_p | max_tokens | error | model |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 83 | TOTAL | 41 models | 259079 | 140807 | 0.5434905955326368 | 25911 | {} | 2h 06m 20s | 1 | 0.0 | 0.95 | 1024 | Models: dsr-one-example-grpo_global_step_50, dsr-one-example-grpo_global_step_100, dsr-one-example-grpo_global_step_150, dsr-one-example-grpo_global_step_200, dsr-one-example-grpo_global_step_250, dsr-one-example-grpo_global_step_300, dsr-one-example-grpo_global_step_350, dsr-one-example-grpo_global_step_400, dsr-one-example-grpo_global_step_450, dsr-one-example-grpo_global_step_500, dsr-one-example-grpo_global_step_550, dsr-one-example-grpo_global_step_600, dsr-one-example-grpo_global_step_650, dsr-one-example-grpo_global_step_700, dsr-one-example-grpo_global_step_750, dsr-one-example-grpo_global_step_800, dsr-one-example-grpo_global_step_850, dsr-one-example-grpo_global_step_900, dsr-one-example-grpo_global_step_950, dsr-one-example-grpo_global_step_1000, grpo-one-example-gsm8k_global_step_50, grpo-one-example-gsm8k_global_step_100, grpo-one-example-gsm8k_global_step_150, grpo-one-example-gsm8k_global_step_200, grpo-one-example-gsm8k_global_step_250, grpo-one-example-gsm8k_global_step_300, grpo-one-example-gsm8k_global_step_350, grpo-one-example-gsm8k_global_step_400, grpo-one-example-gsm8k_global_step_450, grpo-one-example-gsm8k_global_step_500, grpo-one-example-gsm8k_global_step_550, grpo-one-example-gsm8k_global_step_600, grpo-one-example-gsm8k_global_step_650, grpo-one-example-gsm8k_global_step_700, grpo-one-example-gsm8k_global_step_750, grpo-one-example-gsm8k_global_step_800, grpo-one-example-gsm8k_global_step_850, grpo-one-example-gsm8k_global_step_900, grpo-one-example-gsm8k_global_step_950, grpo-one-example-gsm8k_global_step_1000, Qwen_Qwen2.5-1.5B-Instruct | Tasks: gsm8k_main(0), hendrycks_math(0) |