First, thank you for creating this impressive benchmark for evaluating LLMs' ability to solve SQL issues. BIRD-CRITIC-1 addresses a critical real-world application need.
Issue Description
While reviewing the repository, I noticed limited documentation on how the various evaluation metrics (Soft EX, Soft EX + Parsing, Test Case, Query Execution Plan) are calculated in detail. This information would be valuable for researchers attempting to reproduce or build upon your work.
Requested Improvements
- Add technical documentation explaining how each evaluation metric is implemented
- Provide examples of successful vs. unsuccessful solutions for each metric type
- Consider adding an evaluation walkthrough for one or two example tasks
This additional documentation would greatly enhance the accessibility of your benchmark for the research community.
Thank you for your consideration.
First, thank you for creating this impressive benchmark for evaluating LLMs' ability to solve SQL issues. BIRD-CRITIC-1 addresses a critical real-world application need.
Issue Description
While reviewing the repository, I noticed limited documentation on how the various evaluation metrics (Soft EX, Soft EX + Parsing, Test Case, Query Execution Plan) are calculated in detail. This information would be valuable for researchers attempting to reproduce or build upon your work.
Requested Improvements
This additional documentation would greatly enhance the accessibility of your benchmark for the research community.
Thank you for your consideration.