Hyungkeun Park

AI Engineer — Model Compression & Efficient Inference (Quantization)

AI Engineer at Nota AI (Platform Team, Quantizer part). Works on post-training quantization for LLMs and graph-level quantization for deployment across heterogeneous edge backends (NetsPresso Quantizer). Previously a display-IC design/verification engineer at Samsung Display.

Work Experience

Nota AI · AI Engineer — Platform Team, Quantizer PartFeb. 2025 – Sep. 2026
  • Implemented PyTorch eager-mode LLM quantization algorithms (QuaRot, GPTQ) in NetsPresso's Advanced Quantizer.
  • Developed a graph-level quantization framework supporting deployment-ready execution across heterogeneous backends including XNNPACK, Arm, and QNN.
  • Enabled backend-specific graph quantization pipelines for executable inference on edge and hardware-accelerated environments.
  • Optimized customer and demo models for quantized deployment and performance validation.
Samsung Display · EngineerJan. 2019 – Aug. 2022
  • Developed and verified QD-Display sensing-data calibration filtering IP (4.3M gate count).
  • Designed LUT-based sensing-data calibration and defect-data filtering modules.
  • Performed design lint (Spyglass), timing closure (Design Compiler), and simulation-based verification (SystemVerilog, SimVision).
  • Verified the design on FPGA (Arria 10, Stratix 10).

Projects

QuaRot for a Production Quantizer · sole implementer2025
  • Production-grade QuaRot (rotation-based outlier removal) in NetsPresso's Advanced Quantizer, extended beyond the official repo's model coverage.
Graph Quantizer Edge Deployment · backend deployment owner2025–2026
  • Deployment-ready graph quantization across ExecuTorch XNNPACK, Qualcomm QNN, and Arm Ethos-U — constraints as data, upstream fixes over local patches.
Automatic Mixed-Precision Quantization · feature owner2025–2026
  • Per-layer precision search for graph quantization — from the first ratio mode to a shipped CLI scheme and a 3-axis extension architecture.

Open Source Contributions

Hugging Face Transformers · huggingface/transformers
  • Fixed model re-saving error on NVFP4-unsupported devices — #44983.
PyTorch ExecuTorch · pytorch/executorch
Netron · lutzroeder/netron
  • Reduced read amplification when opening large models — #1597 open.

Education

Yonsei University · M.S.Mar. 2023 – Feb. 2025
  • Integrated Technology (School of Integrated Technology)
  • GPA 4.18/4.3 · Supervisor: Jong-Seok Lee (MCML Lab.) · Research interests: knowledge distillation, quantization.
Inha University · B.S.Mar. 2013 – Feb. 2019
  • Information Communication Technology
  • GPA 4.05/4.5 · Merit-based scholarship (2014-1, 2017-1)

Publications

  • Adaptive Explicit Knowledge Transfer for Knowledge Distillation. arXiv:2409.01679, 2024.

Patents

  • Display device
    Application No. 10-2020-0116238 KR / US / CN
  • Display Device Performing Image Sticking Compensation, and Method of Compensating Image Sticking in a Display Device
    Application No. 10-2020-0090210 KR / US / CN