Minsu Kim

Collaboration & Contact

  • Actively looking for collaborators → For non-commercial academic research on LLM agents for chip design, please reach out at minsukim.ai at gmail dot com. I pursue this research independently and would participate in a personal academic capacity, rather than on behalf of my employer.
  • Professional contact → For seminars, workshops, or other engagements related to my professional role, please reach me at minsukim at microsoft dot com.
profile_minsu.jpg

At Microsoft Frontier Tuning, I work on reinforcement learning methods for post-training frontier language models in noisy environments and on long-horizon tasks. Previously, I was a postdoctoral fellow working with Prof. Yoshua Bengio at Mila and Prof. Sungjin Ahn and Prof. Sungsoo Ahn at KAIST, focusing on structured reasoning for trustworthy LLMs.

I received my Ph.D. from KAIST in Prof. Jinkyoo Park’s group, studying reinforcement learning for combinatorial optimization and its applications to LLMs. During my M.S. at KAIST in Prof. Joungho Kim’s group, I studied learning-based physical-layout optimization for semiconductor systems. I received my B.S. in Mathematics and Computer Science from KAIST.

Independent Research Interest

Separately from my professional role, I am independently exploring how LLM agents and reinforcement learning can optimize computing systems and semiconductor physical design—an application area connected to my M.S. research. I am particularly interested in system optimization, floorplanning, routing, and chip placement.

  • From code to hardware: The success of coding agents such as Claude Code and Codex shows that LLMs can reason about and optimize complex software, often written in Python. Beneath that software layer are hardware description languages such as Verilog, followed by the physical implementation of circuits. I see these lower layers of the computing stack as a natural next frontier for LLM agents and reinforcement learning.
  • Long-term vision: I am interested in a self-improving loop between LLMs for chips and chips for LLMs: agents help design more capable and efficient hardware, and that hardware, in turn, enables more capable models and agents.

Research Themes in Industry

In my professional research, I focus on practical methods that help LLMs tackle reasoning and agentic tasks that current models cannot yet handle reliably. I am particularly interested in settings without readily verifiable rewards—unlike many math and coding tasks—or where success depends on decisions over much longer horizons. The following themes summarize my goals and methods:

  • User-Aligned Reasoning Models & Agents: multi-turn reasoning grounded in user intent, multi-agent orchestration, and reward modeling
  • Reinforcement Learning at Scale: long-horizon credit assignment, structured exploration, replay-based training, and sample-efficient learning for LLM post-training

Academic Service

  • Area Chair: NeurIPS (Position Paper Track, 2026)
  • Reviewer (Conferences): NeurIPS (2022–2025), ICML (2023–2026), ICLR (2024–2026)
  • Reviewer (Journals): TNNLS (2025–2026), TPAMI (2025), TMLR (2025)

News

Aug 05, 2026 I joined Microsoft Frontier Tuning as a Senior Research Scientist.
Jul 08, 2026 Our paper, Self-Evolving Curriculum for LLM Reasoning, was accepted to COLM 2026!
May 01, 2026 Two papers—Active Attacks and S3GFN—were accepted to ICML 2026!
Feb 08, 2026 Two papers—LVI and DAV—were accepted to ICLR 2026!
Sep 25, 2025 Four papers—SGDS, TBA, EGM, and ABCD—were accepted to NeurIPS 2025!

Selected Publications

  1. ICML
    Active Attacks: Red-teaming LLMs via Adaptive Environments
    Taeyoung Yun, Pierre-Luc St-Charles, Jinkyoo Park, Yoshua Bengio, and Minsu Kim
    International Conference on Machine Learning, 2026
  2. ICLR
    Latent Veracity Inference for Identifying Errors in Stepwise Reasoning
    Minsu Kim*, Jean-Pierre Falet*, Oliver E Richardson, Xiaoyin Chen, Moksh Jain, Sungjin Ahn, Sungsoo Ahn, and Yoshua Bengio
    International Conference on Learning Representations, 2026
  3. NeurIPS
    On Scalable and Efficient Training of Diffusion Samplers
    Minkyu Kim*, Kiyoung Seong*, Dongyeop Woo, Sungsoo Ahn, and Minsu Kim
    Advances in Neural Information Processing Systems, 2025
  4. Thesis
    Off-policy Training Methods for Probabilistic Agents in Combinatorial Space
    Minsu Kim
    Korea Advanced Institute of Science and Technology (KAIST), 2025
  5. AISTATS
    Ant Colony Sampling with GFlowNets for Combinatorial Optimization
    Minsu Kim*, Sanghyeok Choi*, Jiwoo Son, Hyeonah Kim, Jinkyoo Park, and Yoshua Bengio
    International Conference on Artificial Intelligence and Statistics, 2025
  6. ICLR
    Adaptive Teachers for Amortized Samplers
    Minsu Kim*, Sanghyeok Choi*, Taeyoung Yun, Emmanuel Bengio, Leo Feng, Jarrid Rector-Brooks, Sungsoo Ahn, Jinkyoo Park, Nikolay Malkin, and Yoshua Bengio
    International Conference on Learning Representations, 2025
  7. ICML
    Learning to Scale Logits for Temperature-Conditional GFlowNets
    Minsu Kim*, Joohwan Ko*, Taeyoung Yun*, Dinghuai Zhang, Ling Pan, Woochang Kim, Jinkyoo Park, Emmanuel Bengio, and Yoshua Bengio
    International Conference on Machine Learning, 2024
  8. ICLR
    Local Search GFlowNets
    Minsu Kim, Taeyoung Yun, Emmanuel Bengio, Dinghuai Zhang, Yoshua Bengio, Sungsoo Ahn, and Jinkyoo Park
    International Conference on Learning Representations, 2024
  9. NeurIPS
    Bootstrapped Training of Score-Conditioned Generator for Offline Design of Biological Sequences
    Minsu Kim, Federico Berto, Sungsoo Ahn, and Jinkyoo Park
    Advances in Neural Information Processing Systems, 2023
  10. NeurIPS
    Sym-NCO: Leveraging Symmetricity for Neural Combinatorial Optimization
    Minsu Kim, Junyoung Park, and Jinkyoo Park
    Advances in Neural Information Processing Systems, 2022
  11. NeurIPS
    Learning collaborative policies to solve NP-hard routing problems
    Minsu Kim, Jinkyoo Park, and Joungho Kim
    Advances in Neural Information Processing Systems, 2021