ReZero-Search-LLM-Agent-Fork

Commit Graph

Author	SHA1	Message	Date
thinhlpg	d0e6068055	fix: strengthen reward correctness logic to handle final message is not asnwer form assistant. Also update logs for reward functions for better debug - Added 'logs/' directory to .gitignore to exclude log files. - Introduced log_chat_state function to log chat states and rewards to JSONL files. - Updated reward functions to log chat states with validation results for better tracking and debugging.	8 months ago
thinhlpg	1bd609dfae	test: enhance reward correctness tests with validation logic - Updated test cases to include role and tag validation for assistant messages. - Ensured that only properly formatted messages with answer tags are accepted. - Added new test for validating various incorrect formats and their expected outcomes.	8 months ago
thinhlpg	338655e563	feat: refine user prompt logic for improved clarity and structure	8 months ago
thinhlpg	6d994feeb2	feat: enhance evaluation scripts for base and LoRA models	8 months ago
thinhlpg	da60b52bd1	feat: refactor download and upload scripts for improved argument handling (more notebook friendly :D)	8 months ago
thinhlpg	fa3c0562fe	feat: add evaluation scripts for base and LoRA models - Introduced `eval_base.py` for evaluating base model performance. - Introduced `eval_lora.py` for evaluating LoRA model performance with additional LoRA weight handling.	8 months ago
thinhlpg	1047e2fa1c	chore: update .gitignore and requirements for unsloth versions	8 months ago
thinhlpg	83f86869f6	chore: update .gitignore and add new toys data files	8 months ago
thinhlpg	133cb1ab90	test: add Qwen tokenizer adapter tests Implemented unit tests for the Qwen tokenizer adapter, including format handling, mask generation, and multi-turn conversation support	8 months ago
thinhlpg	6efe01d5ff	chore: update Makefile and requirements for testing - Added a 'test' target in Makefile to run unit tests using pytest. - Included 'wandb' in requirements.txt for experiment tracking.	8 months ago
thinhlpg	af7f38c792	feat: add code for qwen architecture	8 months ago
thinhlpg	e7915a6a8e	feat: add util script to upload/download checkpoints	8 months ago
thinhlpg	9009440663	chore: disable logging, enable torch complie	8 months ago
thinhlpg	d2f03b96ab	feat: enhance evaluation script and remove deprecated shell script - Updated eval.py to streamline model evaluation using vLLM and unsloth. - Deleted eval.sh as its functionality is now integrated into eval.py. - Updated .gitignore to exclude eval_logs directory.	8 months ago
thinhlpg	908768458c	chore: update Makefile and requirements for testing - Added 'tests' directory to check_dirs in Makefile for better organization. - Included 'pytest' in requirements.txt to facilitate unit testing.	8 months ago
thinhlpg	90b45c62ab	docs: update docs and notebooks for the past few days, (observation, debugging) - observation: model hallucniate the search result, docs about debugigng and adapting to r1 distil base model, notebooks on the detail of making training r1 distil works	8 months ago
thinhlpg	3910ef343a	test: add unit tests for agent, reward functions, and tokenizer adapters	8 months ago
thinhlpg	31dcbf5d8a	feat: refactor whole code base, add logic for training R1 distil base models, change some template and reward logics - Break down rl_helpers into smaller modules - Removed deprecated rl_helpers module to streamline the codebase. - Enhance initial user prompt template inspired by Search-R1	8 months ago
thinhlpg	c90c03267e	feat: change user prompt template to search-r1 inspried format use <search></search> instead of embed whole tool definition, which resulted in lots or parsing errors	8 months ago
thinhlpg	58dcf9a99d	refactor: simplify inference script by removing logger, load 16 bit model intead of raw lora finetuned	8 months ago
thinhlpg	da79e986b6	feat: add new script and functionality in train script to save model in 16 bit format	8 months ago
thinhlpg	f6b6cca2ce	feat: add multiple reference notebooks for model training and inference Big thanks to author(s) for the great reference code!	8 months ago
thinhlpg	04593fa8fd	style: change line length to 119, organize imports	8 months ago
thinhlpg	abb18b10d8	feat: add CLI inference script with search functionality This script is a bit dumb, but it worked. I'll update it later XD	8 months ago
thinhlpg	fe70896023	chore: add Makefile for installation, code quality checks, style formatting, cleanup, and other tasks	8 months ago
thinhlpg	60233f2113	chore: update .gitignore	8 months ago
thinhlpg	fd32bcacfd	chores: update worklog and research progress	8 months ago
thinhlpg	37730095a9	feat: add eval scripts that compare base model performance with the grpo trained model	8 months ago
thinhlpg	7f2f43aa46	chore: clean up notebooks	8 months ago
thinhlpg	3c2deaced9	refactor: restructure code base, better centralize logging logic	8 months ago
thinhlpg	04d56325bb	feat: add new reward functions, add less dumb data generation logic, implement better logging	9 months ago
thinhlpg	b22b02ea1d	feat: changed `<reasoning>` tags to `<think>	9 months ago
thinhlpg	7d4de89186	chore: update worklog 250324 - Added `train_autodidact_1B.py` for quick test. - Update `00_worklog.md`, `dataset.md`, and `reward-functions.md` to reflect new training strategies and reward functions.	9 months ago
thinhlpg	1bdee261b6	feat: add draft data generation and documentation - Updated `00_worklog.md` to reflect optimizations for speed and quality in dataset generation. - Introduced new documentation files: `choosing-llm-and-prompt-101.md`, `ds-pipeline-v0.md`, and `paraphrase-prompt.md` for better clarity on LLM choices and dataset pipeline. - Added a Jupyter notebook `250324_generate_data_anatomy.ipynb` to explore the data generation process	9 months ago
thinhlpg	f19354a8c9	chore: clean up notebook output	9 months ago
thinhlpg	f60ab499eb	chore: update worklog	9 months ago
thinhlpg	a58722e16f	feat: add initial project structure and core functionality - Added initial files from AutoDiact as starting point - Enhanced `README.md` with project overview and setup instructions. . - Removed `ugly_code_file.py` as part of cleanup. - Added various documentation files and assets for project clarity. - Included Jupyter notebooks for training and experimentation.	9 months ago
Thinh Le	91c2476c28	chore: initial commit - the ugliest code i've ever written 💀 Dropping this absolute disaster of a code file to break the paralysis. No more overthinking, no more perfectionism—just write, make it work, and refine later. Starting this repo with the most unreadable, unformatted, and ugly code possible. The goal? Trick my brain into not caring about style—just build. This mess exists to remind me that progress > perfection. Ship first, clean up later.	9 months ago
Thinh Le	bf32fdd897	Initial commit	9 months ago

39 Commits (d0e6068055eac86f83c4900c393495103bf294d8) All Branches Search

39 Commits (d0e6068055eac86f83c4900c393495103bf294d8)

All Branches