Reward models used in reinforcement learning from human feedback are vulnerable to reward hacking: as the policy maximizes a learned proxy reward, true quality plateaus or degrades. We assume that reward hacking is often caused by flipped advantage signs: instead of reducing the likelihood of a bad response, a flipped sign causes the update to increase it. By considering an adversarial perturbation in the reward model parameter space, we derive a certified sign-preservation radius, the smallest perturbation that can flip the advantage sign during policy optimization. We propose Sign-Certified Policy Optimization (SignCert-PO), which down-weights non-robust completions in the policy gradient update. Unlike prior approaches that require multiple reward models or access to the reward model training data, SignCert-PO is lightweight and operates purely at the policy optimization stage, using only the reward model parameters and on-policy completions. On TL;DR summarization and AlpacaFarm, SignCert-PO consistently achieves a better win rate than baselines and reduces reward hacking.
@inproceedings{ono2026signcert,title={Mitigating Reward Hacking in {RLHF} via Advantage Sign Robustness},author={Ono, Shinnosuke and Ackermann, Johannes and Nishimori, Soichiro and Ishida, Takashi and Sugiyama, Masashi},booktitle={2nd Workshop on Epistemic Intelligence in Machine Learning (EIML) at ICML},year={2026},month=apr,doi={10.48550/arXiv.2604.02986},}
A Japanese Language Model and Three New Evaluation Benchmarks for Pharmaceutical NLP
Shinnosuke Ono, Issey Sukeda, Takuro Fujii, and 2 more authors
In Proceedings of the 4th International Joint Conference on Natural Language Processing and the 14th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (IJCNLP-AACL)equal contribution with I. Sukeda , Dec 2025
We present a Japanese domain-specific language model for the pharmaceutical field, developed through continual pretraining on 2 billion Japanese pharmaceutical tokens and 8 billion English biomedical tokens. To enable rigorous evaluation, we introduce three new benchmarks: YakugakuQA, based on national pharmacist licensing exams; NayoseQA, which tests cross-lingual synonym and terminology normalization; and SogoCheck, a task designed to assess consistency reasoning between paired statements. We evaluate our model against both open-source medical language models and commercial models, including GPT-4o. Results show that our domain-specific model outperforms existing open models and achieves competitive performance with commercial ones, particularly on terminology-heavy and knowledge-based tasks. Even GPT-4o performs poorly on SogoCheck, suggesting that cross-sentence consistency reasoning remains an open challenge. Our benchmark suite offers a broader diagnostic lens for pharmaceutical NLP, covering factual recall, lexical variation, and logical consistency.
@inproceedings{ono2025japanese,title={A Japanese Language Model and Three New Evaluation Benchmarks for Pharmaceutical {NLP}},author={Ono, Shinnosuke and Sukeda, Issey and Fujii, Takuro and Buma, Kosei and Sasaki, Shunsuke},booktitle={Proceedings of the 4th International Joint Conference on Natural Language Processing and the 14th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (IJCNLP-AACL)},year={2025},month=dec,doi={10.48550/arXiv.2505.16661},}
preprint
AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation
Ryosuke Takanami, Petr Khrapchenkov, Shu Morikuni, and 32 more authors
As robots move from controlled settings to unstructured human environments, building generalist agents that reliably follow natural language instructions remains a central challenge. Progress in robust mobile manipulation requires large-scale multimodal datasets that capture contact-rich and long-horizon tasks, yet existing resources lack synchronized force-torque sensing, hierarchical annotations, and explicit failure cases. We address this gap with the AIRoA MoMa Dataset, a large-scale real-world multimodal dataset for mobile manipulation. It includes synchronized RGB images, joint states, six-axis wrist force-torque signals, and internal robot states, together with a two-layer annotation schema of sub-goals and primitive actions for hierarchical learning and error analysis. The initial dataset comprises 25,469 episodes (about 94 hours) collected with the Human Support Robot and is standardized in the LeRobot v2.1 format.
@article{takanami2025airoa,title={{AIRoA} {MoMa} Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation},author={Takanami, Ryosuke and Khrapchenkov, Petr and Morikuni, Shu and Arima, Jumpei and Takaba, Yuta and Maeda, Shunsuke and Okubo, Takuya and Sano, Genki and Sekioka, Satoshi and Kadoya, Aoi and Kambara, Motonari and Nishiura, Naoya and Suzuki, Haruto and Yoshimoto, Takanori and Sakamoto, Koya and Ono, Shinnosuke and Yang, Hu and Yashima, Daichi and Horo, Aoi and Motoda, Tomohiro and Chiyoma, Kensuke and Ito, Hiroshi and Fukuda, Koki and Goto, Akihito and Morinaga, Kazumi and Ikeda, Yuya and Kawada, Riko and Yoshikawa, Masaki and Kosuge, Norio and Noguchi, Yuki and Ota, Kei and Matsushima, Tatsuya and Iwasawa, Yusuke and Matsuo, Yutaka and Ogata, Tetsuya},journal={arXiv preprint arXiv:2509.25032},year={2025},month=sep,doi={10.48550/arXiv.2509.25032},}