Hyungkeun Park
AI Engineer — Model Compression & Efficient Inference (Quantization)
AI Engineer at Nota AI (Platform Team, Quantizer part). Works on post-training quantization for LLMs and graph-level quantization for deployment across heterogeneous edge backends (NetsPresso Quantizer). Previously a display-IC design/verification engineer at Samsung Display.
Work Experience
Nota AI · AI Engineer — Platform Team, Quantizer PartFeb. 2025 – Sep. 2026
- Implemented PyTorch eager-mode LLM quantization algorithms (QuaRot, GPTQ) in NetsPresso's Advanced Quantizer.
- Developed a graph-level quantization framework supporting deployment-ready execution across heterogeneous backends including XNNPACK, Arm, and QNN.
- Enabled backend-specific graph quantization pipelines for executable inference on edge and hardware-accelerated environments.
- Optimized customer and demo models for quantized deployment and performance validation.
Samsung Display · EngineerJan. 2019 – Aug. 2022
- Developed and verified QD-Display sensing-data calibration filtering IP (4.3M gate count).
- Designed LUT-based sensing-data calibration and defect-data filtering modules.
- Performed design lint (Spyglass), timing closure (Design Compiler), and simulation-based verification (SystemVerilog, SimVision).
- Verified the design on FPGA (Arria 10, Stratix 10).
Projects
QuaRot for a Production Quantizer · sole implementer2025
- Production-grade QuaRot (rotation-based outlier removal) in NetsPresso's Advanced Quantizer, extended beyond the official repo's model coverage.
Graph Quantizer Edge Deployment · backend deployment owner2025–2026
- Deployment-ready graph quantization across ExecuTorch XNNPACK, Qualcomm QNN, and Arm Ethos-U — constraints as data, upstream fixes over local patches.
Automatic Mixed-Precision Quantization · feature owner2025–2026
- Per-layer precision search for graph quantization — from the first ratio mode to a shipped CLI scheme and a 3-axis extension architecture.
Open Source Contributions
Hugging Face Transformers · huggingface/transformers
- Fixed model re-saving error on NVFP4-unsupported devices — #44983.
PyTorch ExecuTorch · pytorch/executorch
Netron · lutzroeder/netron
- Reduced read amplification when opening large models — #1597 open.
Education
Yonsei University · M.S.Mar. 2023 – Feb. 2025
- Integrated Technology (School of Integrated Technology)
- GPA 4.18/4.3 · Supervisor: Jong-Seok Lee (MCML Lab.) · Research interests: knowledge distillation, quantization.
Inha University · B.S.Mar. 2013 – Feb. 2019
- Information Communication Technology
- GPA 4.05/4.5 · Merit-based scholarship (2014-1, 2017-1)
Publications
- Adaptive Explicit Knowledge Transfer for Knowledge Distillation. arXiv:2409.01679, 2024.
Patents
- Display device
Application No. 10-2020-0116238 KR / US / CN - Display Device Performing Image Sticking Compensation, and Method of Compensating Image Sticking in a Display Device
Application No. 10-2020-0090210 KR / US / CN