3 Commits (31dcbf5d8af3dd6fb2ee0eb026c62b48a282e03d)

Author SHA1 Message Date
thinhlpg c90c03267e feat: change user prompt template to search-r1 inspried format
2 months ago
thinhlpg 04593fa8fd style: change line length to 119, organize imports
2 months ago
thinhlpg 37730095a9 feat: add eval scripts that compare base model performance with the grpo trained model
2 months ago