Publications
2026
ReSET: Accurate Latency-Critical NVFP4 Reasoning via Step-Aware Temperature Scaling
Sihwa Lee*, Janghwan Lee*, Donghoon Yoo, Jae Gon Kim, Hanyul Ryu, Soojung Ryu, Jungwook Choi
Preprint, 2026
ReQAT: Achieving Full-Precision Reasoning Accuracy with 4-bit Floating-Point Quantization-Aware Training
Janghwan Lee, Sihwa Lee, Jinseok Kim, Yongjik Kim, Jieun Lim, Jinwook Oh, Jungwook Choi
Forty-third International Conference on Machine Learning (ICML, Oral), 2026
2025
AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference
Janghwan Lee, Jiwoong Park, Jinseok Kim, Yongjik Kim, Jungju Oh, Jinwook Oh, Jungwook Choi
Findings of the Association for Computational Linguistics (ACL Findings), 2025
RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy
Geonho Lee*, Janghwan Lee*, Sukjin Hong*, Minsoo Kim, Euijai Ahn, Du-Seong Chang, Jungwook Choi
The 39th Annual AAAI Conference on Artificial Intelligence (AAAI), 2025
2024
Improving Conversational Abilities of Quantized Large Language Models via Direct Preference Alignment
Janghwan Lee*, Seongmin Park*, Sukjin Hong, Minsoo Kim, Du-Seong Chang, Jungwook Choi
The 62nd Annual Meeting of the Association for Computational Linguistics (ACL, Oral), 2024
ISP2DLA: Automated Deep Learning Accelerator Design for On-Sensor Image Signal Processing
Dong-eon Won*, Yeeun Kim*, Janghwan Lee, Minjae Lee, Jonghyun Bae, Jongjoo Park, Jeongyong Song, Jungwook Choi
35th IEEE International Conference on Application-specific Systems, Architectures and Processors (ASAP, Poster), 2024
Searching Optimal Floating-Point Format for Sub-8-Bit Large Language Model Inference
Youngdeok Hwang*, Janghwan Lee*, Jiwoong Park, Jieun Lim, Jungwook Choi
International Conference on Electronics, Information, and Communication (ICEIC, Oral), 2024
SPADE: Sparse Pillar-based 3D Object Detection Accelerator for Autonomous Driving
Minjae Lee, Seongmin Park, Hyungmin Kim, Minyong Yoon, Janghwan Lee, Junwon Choi, Nam Sung Kim, Mingu Kang, Jungwook Choi
30th IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2024
2023
Enhancing Computation Efficiency in Large Language Models through Weight and Activation Quantization
Janghwan Lee*, Minsoo Kim*, Seungcheol Baek, Seokjoong Hwang, Wonyong Sung, Jungwook Choi
The 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023
Token-Scaled Logit Distillation for Ternary Weight Generative Language Models
Minsoo Kim, Sihwa Lee, Janghwan Lee, Sukjin Hong, Duseong Chang, Wonyong Sung, Jungwook Choi
Thirty-seventh Conference on Neural Information Processing Systems (NeurIPS), 2023
Range-Invariant Approximation of Non-Linear Operations for Efficient BERT Fine-Tuning
Janghyeon Kim, Janghwan Lee, JeongHo Han, Sangheon Lee, Jungwook Choi
2023 60th ACM/IEEE Design Automation Conference (DAC), 2023
Finding Optimal Numerical Format for Sub-8-Bit Post-Training Quantization of Vision Transformers
Janghwan Lee, Youngdeok Hwang, Jungwook Choi
2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023
2022
Optimizing Exponent Bias for Sub-8bit Floating-Point Inference of Fine-tuned Transformers
Janghwan Lee, Jungwook Choi
2022 IEEE 4th International Conference on Artificial Intelligence Circuits and Systems (AICAS, Oral), 2022