Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning
This study addresses the performance-safety imbalance in enterprise-level agent tool calling by proposing an Anchored Supervised Fine-Tuning strategy. Built upon a Mixture-of-Experts architecture, this approach integrates a hybrid Muon-Adam optimizer with Kullback-Leibler divergence constraints, leveraging synthetic trajectory data to achieve stable alignment under few-shot conditions while enhancing tool-use capabilities and ensuring safety. Experimental results demonstrate that the model achieves a score of 0.785 on BFCL Core and attains the highest average across six benchmarks. Notably, it significantly outperforms its predecessor in Writer Agent tasks while maintaining superior performance in safety and bias evaluations. Collectively, these findings establish a novel paradigm for the efficient and secure deployment of enterprise intelligent agents.