Regarding the prompt used for evaluating the DeepSeek-8B model on the GPQA-D dataset, I noticed while reading the paper that you did not specify exactly what prompt was used for GPQA-D on DeepSeek-8B. Could you provide the exact prompt? I found that when I use the default prompt to test DeepSeek-8B on GPQA-D, the results differ significantly from the numbers reported in the paper. Interestingly, using the same prompt to test the Qwen3-32B model on GPQA-D does not cause this issue.
Regarding the prompt used for evaluating the DeepSeek-8B model on the GPQA-D dataset, I noticed while reading the paper that you did not specify exactly what prompt was used for GPQA-D on DeepSeek-8B. Could you provide the exact prompt? I found that when I use the default prompt to test DeepSeek-8B on GPQA-D, the results differ significantly from the numbers reported in the paper. Interestingly, using the same prompt to test the Qwen3-32B model on GPQA-D does not cause this issue.