Trained with DPO for 100 steps on an mixed synthetic (K3 and R1+R1-0528) writing dataset inspired from various private data, and human vs synthetic.
Negatives were built with high-slop models and middle-of-the-pack writing models for the synthetic set, and middle-of-the-pack writing models for human vs synthetic.
About a fifth of the rows included trace inversion from Iris-12B-v1.3.2 on the synthetic data (to better fit to her own reasoning traces).