Skip to content
Navigation Menu
Sign in
Appearance settings
Platform
AI CODE CREATION
GitHub Copilot
Write better code with AI
GitHub Copilot app
Direct agents from issue to merge
MCP Registry
Integrate external tools
DEVELOPER WORKFLOWS
Actions
Automate any workflow
Codespaces
Instant dev environments
Issues
Plan and track work
Code Review
Manage code changes
Code Quality
Enforce quality at merge
APPLICATION SECURITY
GitHub Advanced Security
Find and fix vulnerabilities
Code security
Secure your code as you build
Secret protection
Stop leaks before they start
EXPLORE
Why GitHub
Documentation
Blog
Changelog
Marketplace
View all features
Solutions
BY COMPANY SIZE
Enterprises
Small and medium teams
Startups
Nonprofits
BY USE CASE
App Modernization
DevSecOps
DevOps
CI/CD
View all use cases
BY INDUSTRY
Healthcare
Financial services
Manufacturing
Government
View all industries
View all solutions
Resources
EXPLORE BY TOPIC
AI
Software Development
DevOps
Security
View all topics
EXPLORE BY TYPE
Customer stories
Events & webinars
Ebooks & reports
Business insights
GitHub Skills
SUPPORT & SERVICES
Documentation
Customer support
Community forum
Trust center
Partners
View all resources
Open Source
COMMUNITY
GitHub Sponsors
Fund open source developers
PROGRAMS
Security Lab
Maintainer Community
GitHub Stars
Archive Program
REPOSITORIES
Topics
Trending
Collections
Enterprise
ENTERPRISE SOLUTIONS
Enterprise platform
AI-powered developer platform
AVAILABLE ADD-ONS
GitHub Advanced Security
Enterprise-grade security features
Copilot for Business
Enterprise-grade AI features
Premium Support
Enterprise-grade 24/7 support
Pricing
Search
/
Sign in
Sign up
Appearance settings
You signed in with another tab or window.
Reload
to refresh your session.
You signed out in another tab or window.
Reload
to refresh your session.
You switched accounts on another tab or window.
Reload
to refresh your session.
Dismiss alert
{{ message }}
Kav-K
/
agent-challenge
Public
Notifications
You must be signed in to change notification settings
Fork
1
Star
4
Code
Issues
0
Pull requests
0
Actions
Projects
Security and quality
0
Insights
Additional navigation options
Code
Issues
Pull requests
Actions
Projects
Security and quality
Insights
Actions: Kav-K/agent-challenge
Actions
All workflows
Workflows
Publish to PyPI and npm
Publish to PyPI and npm
Tests
Tests
Show more workflows...
Management
Caches
All workflows
All workflows
Actions
Loading...
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading.
Please reload this page
.
will be ignored since log searching is not yet available
Showing runs from all workflows
will be ignored since log searching is not yet available
37 workflow runs
37 workflow runs
Workflow
Filter by Workflow
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching workflows.
Event
Filter by Event
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching events.
Status
Filter by Status
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching statuses.
Branch
Filter by Branch
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching branches.
Actor
Filter by Actor
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching users.
v1.3.0
Publish to PyPI and npm
#6:
Release
v1.3.0
published by
Kav-K
1m 42s
1m 42s
View workflow file
v1.3.0 — Polished chained_arithmetic, expanded template variety
Tests
#31:
Commit
5bc88f6
pushed by
Kav-K
1m 23s
main
main
1m 23s
View workflow file
Remove knowledge facts from chained_arithmetic — GPT-5.2 drops to 90%
Tests
#30:
Commit
7cff50a
pushed by
Kav-K
1m 18s
main
main
1m 18s
View workflow file
Enhance chained_arithmetic with world-knowledge patterns
Tests
#29:
Commit
bc8bb00
pushed by
Kav-K
1m 15s
main
main
1m 15s
View workflow file
Add knowledge_math type (hard tier) — world-knowledge + arithmetic
Tests
#28:
Commit
d895229
pushed by
Kav-K
1m 24s
main
main
1m 24s
View workflow file
v1.2.0 - Stable release
Publish to PyPI and npm
#5:
Release
v1.2.0
published by
Kav-K
37s
37s
View workflow file
v1.2.0 — Empirical tier restructure, human-proof agentic challenges
Tests
#27:
Commit
c1bd5e3
pushed by
Kav-K
1m 13s
main
main
1m 13s
View workflow file
Agentic: remove human-trivial types, expand chained_arithmetic
Tests
#26:
Commit
d9758aa
pushed by
Kav-K
1m 25s
main
main
1m 25s
View workflow file
Add chained_arithmetic (agentic) + power_mod (hard) types
Tests
#25:
Commit
f9dae3f
pushed by
Kav-K
1m 23s
main
main
1m 23s
View workflow file
Relax CI thresholds for new tier structure
Tests
#24:
Commit
81f6913
pushed by
Kav-K
1m 4s
main
main
1m 4s
View workflow file
Fix CI: dynamic mode test uses gpt-4o solver
Tests
#23:
Commit
6e718a8
pushed by
Kav-K
1m 3s
main
main
1m 3s
View workflow file
v1.1.0: Tier restructure — GPT-5.2 100% on all tiers
Tests
#22:
Commit
a25305a
pushed by
Kav-K
1m 20s
main
main
1m 20s
View workflow file
Dynamic mode: iterated prompt for 100% GPT-5.2 solve rate
Tests
#21:
Commit
de41606
pushed by
Kav-K
1m 24s
main
main
1m 24s
View workflow file
GPT-5.2 calibration data + frontend updates
Tests
#20:
Commit
2972ae4
pushed by
Kav-K
1m 26s
main
main
1m 26s
View workflow file
Add 14 safe_solve exact-answer enforcement tests (217 total)
Tests
#19:
Commit
78cfd4b
pushed by
Kav-K
1m 23s
main
main
1m 23s
View workflow file
v1.1.0 — Empirical tier reclassification, agentic dynamic mode, safe_…
Tests
#18:
Commit
ba07d08
pushed by
Kav-K
1m 26s
main
main
1m 26s
View workflow file
safe_solve() exact-answer enforcement + calibration data
Tests
#17:
Commit
af089a3
pushed by
Kav-K
1m 15s
main
main
1m 15s
View workflow file
Make wrong-answer retry test resilient (3 retries)
Tests
#16:
Commit
9a79d5d
pushed by
Kav-K
1m 31s
main
main
1m 31s
View workflow file
Lower medium bulk threshold to 75% for CI variance
Tests
#15:
Commit
1053fd8
pushed by
Kav-K
1m 30s
main
main
1m 30s
View workflow file
Empirical difficulty reclassification with dual-model validation
Tests
#14:
Commit
e33bc37
pushed by
Kav-K
1m 26s
main
main
1m 26s
View workflow file
Adjust easy threshold to 80% (16/20) for CI variance
Tests
#13:
Commit
e468e8a
pushed by
Kav-K
1m 19s
main
main
1m 19s
View workflow file
Reclassify difficulty tiers based on empirical gpt-4o-mini accuracy
Tests
#12:
Commit
c174193
pushed by
Kav-K
1m 23s
main
main
1m 23s
View workflow file
v1.0.0 - Initial release
Publish to PyPI and npm
#4:
Release
v1.0.0
published by
Kav-K
35s
35s
View workflow file
v1.0.0 — Production release
Tests
#11:
Commit
1dd1ff5
pushed by
Kav-K
1m 15s
main
main
1m 15s
View workflow file
Fix live test timeout: increase to 45s, wrap bulk solves in try/except
Tests
#10:
Commit
edf8b07
pushed by
Kav-K
1m 11s
main
main
1m 11s
View workflow file
Previous
1
2
Next
You can’t perform that action at this time.