Official repository for the paper "LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code"
-
Updated
Jul 16, 2025 - Python
Official repository for the paper "LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code"
Evidence-grounded AI agent for Java code repair with LLM-guided patches and human-in-the-loop review.
Exploring and improving the quality of ChatGPT-generated code for LeetCode programming tasks.
基于 AI Agent 服务自动化修复系统:Agent 自动读取错误日志,定位 Bug,生成补丁,运行测试,提交 PR,并通知开发者 Review。AI-powered auto-fix agent for web services: analyzes logs, patches code, runs tests, and creates pull requests automatically.
Repository-level automated code repair agent using SWE-Bench dataset
Trusted autonomy T&E runtime that links mission needs, hazards, scenarios, telemetry, evidence, verification reports, and hash-chained ledgers so AI/autonomous decisions can be reviewed instead of merely trusted.
A reliability layer for AI-built systems: detect failures (tests or runtime drift), reproduce, repair one ticket at a time behind an approval gate, and prove the fix. The safety boundaries most AI agents skip.
Gymnasium RL environment for training LLM agents to autonomously debug and fix Python code with secure sandboxing and test-driven feedback.
🦑 CT 11 — Secure Self-Learning Repair Agent. Droste Fusion. 4 Engines. Reflection Engine. 9/10 Benchmark.
Run broken Python code → it fixes itself. Local-first Python runtime repair.
AI proposes. Humans decide. Source-available AI assurance/control plane for governed code change: agent identity, scoped authorization, policy gates, PR/CI evidence binding, replayable evidence bundles, chained receipts, traceability, and human review.
PyPatch— OpenEnv RL environment where AI agents debug & fix buggy Python code across 3 difficulty levels.
When Free Executors Cost More: The Free-Executor Paradox in Iterative LLM Code-Repair Loops (paper + reproducibility kit)
OpenEnv-based reinforcement learning environment for automated Python code debugging and repair.
5-agent LangGraph system that autonomously resolves GitHub issues, benchmarked across single-agent baseline, paid API, and open-weight HPC configurations on SWE-bench Lite
A safety-constrained coding agent for FastAPI defect diagnosis, isolated repair, revision-aware validation, human approval, and rollback.
Broken code indentation repair tool with automatic language detection
Add a description, image, and links to the code-repair topic page so that developers can more easily learn about it.
To associate your repository with the code-repair topic, visit your repo's landing page and select "manage topics."