Skip to content

Latest commit

 

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 
 
 

README.md

harbor-rl

RL training on Harbor tasks using the harbor.rl interface and Tinker.

harbor.rl turns any Harbor task into an RL environment with step()/grade(). Tinker handles sampling, gradient computation, and weight updates. train.py bridges the two in ~120 lines.

Quick start

uv run harbor_cookbook/harbor_rl/train.py \
    --model moonshotai/Kimi-K2-Thinking \
    --group-size 4 \
    --batch-size 8

Prerequisites