RL training on Harbor tasks using the harbor.rl interface and Tinker.
harbor.rl turns any Harbor task into an RL environment with step()/grade(). Tinker handles sampling, gradient computation, and weight updates. train.py bridges the two in ~120 lines.
uv run harbor_cookbook/harbor_rl/train.py \
--model moonshotai/Kimi-K2-Thinking \
--group-size 4 \
--batch-size 8