Publications

2026

Preprint

ReSET: Accurate Latency-Critical NVFP4 Reasoning via Step-Aware Temperature Scaling

Sihwa Lee*, Janghwan Lee*, Donghoon Yoo, Jae Gon Kim, Hanyul Ryu, Soojung Ryu, Jungwook Choi

Preprint, 2026

PDF
ICML 2026 Oral

ReQAT: Achieving Full-Precision Reasoning Accuracy with 4-bit Floating-Point Quantization-Aware Training

Janghwan Lee, Sihwa Lee, Jinseok Kim, Yongjik Kim, Jieun Lim, Jinwook Oh, Jungwook Choi

Forty-third International Conference on Machine Learning (ICML, Oral), 2026

PDF

2025

ACL 2025 Findings

AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference

Janghwan Lee, Jiwoong Park, Jinseok Kim, Yongjik Kim, Jungju Oh, Jinwook Oh, Jungwook Choi

Findings of the Association for Computational Linguistics (ACL Findings), 2025

PDF
AAAI 2025

RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy

Geonho Lee*, Janghwan Lee*, Sukjin Hong*, Minsoo Kim, Euijai Ahn, Du-Seong Chang, Jungwook Choi

The 39th Annual AAAI Conference on Artificial Intelligence (AAAI), 2025

PDF

2024

ACL 2024 Oral

Improving Conversational Abilities of Quantized Large Language Models via Direct Preference Alignment

Janghwan Lee*, Seongmin Park*, Sukjin Hong, Minsoo Kim, Du-Seong Chang, Jungwook Choi

The 62nd Annual Meeting of the Association for Computational Linguistics (ACL, Oral), 2024

PDF
ASAP 2024

ISP2DLA: Automated Deep Learning Accelerator Design for On-Sensor Image Signal Processing

Dong-eon Won*, Yeeun Kim*, Janghwan Lee, Minjae Lee, Jonghyun Bae, Jongjoo Park, Jeongyong Song, Jungwook Choi

35th IEEE International Conference on Application-specific Systems, Architectures and Processors (ASAP, Poster), 2024

PDF
ICEIC 2024 Oral

Searching Optimal Floating-Point Format for Sub-8-Bit Large Language Model Inference

Youngdeok Hwang*, Janghwan Lee*, Jiwoong Park, Jieun Lim, Jungwook Choi

International Conference on Electronics, Information, and Communication (ICEIC, Oral), 2024

PDF
HPCA 2024

SPADE: Sparse Pillar-based 3D Object Detection Accelerator for Autonomous Driving

Minjae Lee, Seongmin Park, Hyungmin Kim, Minyong Yoon, Janghwan Lee, Junwon Choi, Nam Sung Kim, Mingu Kang, Jungwook Choi

30th IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2024

PDF

2023

EMNLP 2023

Enhancing Computation Efficiency in Large Language Models through Weight and Activation Quantization

Janghwan Lee*, Minsoo Kim*, Seungcheol Baek, Seokjoong Hwang, Wonyong Sung, Jungwook Choi

The 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023

PDF
NeurIPS 2023

Token-Scaled Logit Distillation for Ternary Weight Generative Language Models

Minsoo Kim, Sihwa Lee, Janghwan Lee, Sukjin Hong, Duseong Chang, Wonyong Sung, Jungwook Choi

Thirty-seventh Conference on Neural Information Processing Systems (NeurIPS), 2023

PDF
DAC 2023

Range-Invariant Approximation of Non-Linear Operations for Efficient BERT Fine-Tuning

Janghyeon Kim, Janghwan Lee, JeongHo Han, Sangheon Lee, Jungwook Choi

2023 60th ACM/IEEE Design Automation Conference (DAC), 2023

PDF
ICASSP 2023

Finding Optimal Numerical Format for Sub-8-Bit Post-Training Quantization of Vision Transformers

Janghwan Lee, Youngdeok Hwang, Jungwook Choi

2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023

PDF

2022

AICAS 2022 Oral

Optimizing Exponent Bias for Sub-8bit Floating-Point Inference of Fine-tuned Transformers

Janghwan Lee, Jungwook Choi

2022 IEEE 4th International Conference on Artificial Intelligence Circuits and Systems (AICAS, Oral), 2022

PDF