12 Commits (20fb6779c35d13c6894a1e8bc933d13a474e8e31)

Author SHA1 Message Date
thinhlpg bac5f3b4f7 feat: update config and paths, update data genenration script
2 months ago
thinhlpg 2df9f39fda feat: update model configuration (longer context) and dataset loading logic for improved performance and flexibility
3 months ago
thinhlpg eebf914a81 refactor: moved modules from src/deepsearch to src/
3 months ago
thinhlpg 2fec4f2f42 refactor: change repo stucture (move code from src/ to src/deepsearch)
3 months ago
thinhlpg 1a18cd7bfd feat: update training and evaluation configurations (editable agent generation scripts)
3 months ago
thinhlpg bf480574a2 fix: minor bug
3 months ago
thinhlpg 4de31e0f30 feat: expand reward functions with new strategies and diversity checks
3 months ago
thinhlpg af7f38c792 feat: add code for qwen architecture
3 months ago
thinhlpg 31dcbf5d8a feat: refactor whole code base, add logic for training R1 distil base models, change some template and reward logics
3 months ago
thinhlpg da79e986b6 feat: add new script and functionality in train script to save model in 16 bit format
3 months ago
thinhlpg 04593fa8fd style: change line length to 119, organize imports
3 months ago
thinhlpg 3c2deaced9 refactor: restructure code base, better centralize logging logic
3 months ago