publications

国際学会(査読付き)

2026

  1. codit.png
    Synthesizing Instruction-Tuning Datasets with Contrastive Decoding
    Tatsuya Ichinose, Youmi Ma, Masanari Oi, and 2 more authors
    In COLM, 2026
  2. hatch.jpg
    From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models
    Masanari Oi, Koki Maeda, Ryuto Koike, and 3 more authors
    In ICML, 2026
    Proposed a training framework for vision‑language models to improve multi‑image spatial reasoning
  3. adpo.png
    Autoregressive Direct Preference Optimization
    Masanari Oi, Mahiro Ukai, Masahiro Kaneko, and 2 more authors
    In ICML, 2026
    Identified a theoretical issue in DPO and proposed a new method, ADPO.
  4. mmhe.png
    Multi-modal, Multi-task, Multi-criteria Automatic Evaluation Using a Vision Language Model
    Masanari Oi, Masahiro Kaneko, Naoaki Okazaki, and 1 more author
    In LREC, 2026
  5. swallow_code.png
    Rewriting Pre-Training Data Boosts LLM Performance in Math and Code
    Kazuki Fujii, Yukito Tajima, Sakae Mizuki, and 14 more authors
    In ICLR 2026, 2026
    Proposed a rewriting‑based pipeline for synthesizing mathematics and coding data for LLMs.
  6. discode.png
    DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning
    Nakamasa Inoue, Kanoko Goto, Masanari Oi, and 4 more authors
    In AAAI 2026, 2026

2025

  1. building.png
    Building Instruction-Tuning Datasets from Human-Written Instructions with Open-Weight Large Language Models
    Youmi Ma, Sakae Mizuki, Kazuki Fujii, and 12 more authors
    In COLM 2025, 2025
  2. hall_e.png
    HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
    Yuto Nishimura, Takumi Hirose, Masanari Ohi, and 2 more authors
    In ICLR 2025, 2025
    Proposed an LLM‑based text‑to‑speech model for minute‑long zero-shot synthesis.

2024

  1. cpt.png
    Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities
    Kazuki Fujii, Taishi Nakamura, Mengsay Loem, and 7 more authors
    In COLM 2024, 2024
  2. swallow_corpus.png
    Building a Large Japanese Web Corpus for Large Language Models
    Naoaki Okazaki, Kakeru Hattori, Shota Hirai, and 7 more authors
    In COLM 2024, 2024
  3. likelihood_bias.png
    Likelihood-based Mitigation of Evaluation Bias in Large Language Models
    Masanari Ohi, Masahiro Kaneko, Ryuto Koike, and 2 more authors
    In ACL Findings, 2024
    Detected and mitigated self‑preference bias in LLM‑as‑a‑judge evaluation.

論文誌(査読付き)

2025

  1. 大規模言語モデルにおける評価バイアスの尤度に基づく緩和 (English title: Likelihood-based Mitigation of Evaluation Bias in Large Language Models)
    Masanari Ohi, Masahiro Kaneko, Ryuto Koike, and 2 more authors
    Journal of Natural Language Processing, Jun 2025
    最優秀論文賞.

2024

  1. elp_adapters.png
    ELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing Tasks
    Nakamasa Inoue, Shinta Otake, Takumi Hirose, and 2 more authors
    IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024

ワークショップ(査読付き・非アーカイバル)

2025

  1. why_jp.png
    Why We Build Local Large Language Models: An Observational Analysis from 35 Japanese and Multilingual LLMs
    Koshiro Saito, Sakae Mizuki, Masanari Ohi, and 11 more authors
    In MELT 2025, 2025

プレプリント(査読なし)

現在掲載しているプレプリントはありません。

国内学会・研究会(査読なし)

2026

  1. 自己回帰性を組み込んだ直接選好最適化 (English title: Autoregressive Direct Preference Optimization)
    Masanari Oi, Masahiro Kaneko, Naoki Okazaki, and 1 more author
    In The 32nd Annual Meeting of the Association for Natural Language Processing, 2026
    最優秀賞.
  2. 対照的デコーディングを用いた指示学習データの合成 (English title: Synthesizing Instruction-Tuning Datasets with Contrastive Decoding)
    Tatsuya Ichinose, Youmi Ma, Masanari Oi, and 2 more authors
    In The 32nd Annual Meeting of the Association for Natural Language Processing, 2026
    In Japanese.
  3. 蒸留による日英推論型大規模言語モデル構築戦略の探索 (English title: Exploring Strategies for Building Japanese-English Reasoning Large Language Models through Distillation)
    Sakae Mizuki, Kazuki Fujii, Masaki Kawamura, and 13 more authors
    In The 32nd Annual Meeting of the Association for Natural Language Processing, 2026
    In Japanese.

2025

  1. JUBAKU: 日本文化における偏見評価のための敵対的ベンチマーク (English title: JUBAKU: An Adversarial Benchmark for Evaluating Bias in Japanese Culture)
    Taihei Shiotani, Masahiro Kaneko, Ayana Niwa, and 4 more authors
    In The 39th Annual Conference of the Japanese Society for Artificial Intelligence, 2025
    全国大会優秀賞.
  2. 複数タスク・複数項目に跨ったマルチモーダル自動評価手法 (English title: Multi-modal, Multi-task, Multi-criteria Automatic Evaluation)
    Masanari Oi, Masahiro Kaneko, Naoki Okazaki, and 1 more author
    In The 31st Annual Meeting of the Association for Natural Language Processing, 2025
    委員特別賞.
  3. 模倣学習による大規模言語モデルの指示チューニング (English title: Instruction Tuning for Large Language Models through Imitation Learning)
    Youmi Ma, Sakae Mizuki, Kazuki Fujii, and 12 more authors
    In The 31st Annual Meeting of the Association for Natural Language Processing, 2025
    In Japanese.
  4. 新聞記事からつくる時事と社会に強い日本語LLM (English title: Building a Japanese LLM Strong in Current Affairs and Society from Newspaper Articles)
    Kakeru Hattori, Sakae Mizuki, Kazuki Fujii, and 15 more authors
    In The 31st Annual Meeting of the Association for Natural Language Processing, 2025
    In Japanese.
  5. Swallowコーパスv2: 教育的な日本語ウェブコーパスの構築 (English title: Swallow Corpus v2: Building an Educational Japanese Web Corpus)
    Kakeru Hattori, Naoaki Okazaki, Sakae Mizuki, and 11 more authors
    In The 31st Annual Meeting of the Association for Natural Language Processing, 2025
    In Japanese.

2024

  1. LLMに日本語テキストを学習させる意義 (English title: The Value of Training LLMs on Japanese Text)
    Koshiro Saito, Sakae Mizuki, Masanari Ohi, and 11 more authors
    In The 261st NL Research Presentation, 2024
    優秀研究賞.
  2. 大規模言語モデルにおける評価バイアスの尤度に基づく緩和 (English title: Likelihood-based Mitigation of Evaluation Bias in Large Language Models)
    Masanari Ohi, Masahiro Kaneko, Ryuto Koike, and 2 more authors
    In The 30th Annual Meeting of the Association for Natural Language Processing, 2024
    若手奨励賞.
  3. Swallowコーパス: 日本語大規模ウェブコーパスの構築 (English title: Swallow Corpus: A Large-Scale Japanese Web Corpus)
    Naoaki Okazaki, Kakeru Hattori, Shota Hirai, and 7 more authors
    In The 30th Annual Meeting of the Association for Natural Language Processing, 2024
    優秀賞.
  4. 大規模言語モデルの日本語能力の効率的な強化: 継続事前学習における語彙拡張と対訳コーパスの活用 (English title: Efficiently Enhancing the Japanese Capabilities of Large Language Models: Vocabulary Expansion and Parallel Corpora in Continual Pre-training)
    Sakae Mizuki, Hiroki Iida, Kazuki Fujii, and 7 more authors
    In The 30th Annual Meeting of the Association for Natural Language Processing, 2024
    In Japanese.
  5. 継続事前学習による日本語に強い大規模言語モデルの構築 (English title: Building a Japanese-Strong Large Language Model through Continual Pre-training)
    Kazuki Fujii, Taishi Nakamura, Mengsay Loem, and 7 more authors
    In The 30th Annual Meeting of the Association for Natural Language Processing, 2024
    優秀賞.