Pulse
0
Velocity
Stars
0
A benchmark dataset for evaluating reasoning in large language models.
A benchmark dataset to evaluate reasoning capabilities of LLMs.
Benchmark and advance LLM reasoning with a comprehensive, structured dataset.
Why Trending
Provides standardized evaluation for complex reasoning in LLMs.
Target Audience
AI researchers evaluating LLM reasoning performance.
Similar Projects