6 Commits (3081d6e36b22baea937801fd66031d5e8e85c10f)

Author SHA1 Message Date
thinhlpg 4de31e0f30 feat: expand reward functions with new strategies and diversity checks
1 month ago
thinhlpg af7f38c792 feat: add code for qwen architecture
1 month ago
thinhlpg 31dcbf5d8a feat: refactor whole code base, add logic for training R1 distil base models, change some template and reward logics
1 month ago
thinhlpg da79e986b6 feat: add new script and functionality in train script to save model in 16 bit format
1 month ago
thinhlpg 04593fa8fd style: change line length to 119, organize imports
1 month ago
thinhlpg 3c2deaced9 refactor: restructure code base, better centralize logging logic
1 month ago