Files
2026-09-04 14:58:42 +08:00

1023 BLFS

schema_version, metadata, verifier, agent, environment
schema_version metadata verifier agent environment
1.3
author_name author_email difficulty category subcategory category_confidence task_type modality interface skill_type tags
Jiajun Bao jiajunb@alumni.cmu.edu hard software-engineering debugging high
debugging
repair
source-code
terminal
python
debugging-heuristic
library-api-usage
post-training
trl
debugging
reasoning-model
grpo
type timeout_sec service hardening
test-script 120.0 main
cleanup_conftests
true
timeout_sec
3600.0
network_mode build_timeout_sec os cpus memory_mb storage_mb gpus
public 2400.0 linux 4 2048 10240 0

debug-trl-grpo

I am training a countdown math task using TRL for my model with GRPO, but the new model shows no improvement. The source code is at /app/trl. Check if there is any bug and fix them.

DO NOT modify the training script (/app/train_grpo.py) or reward function (/app/reward_fn.py)