Publications

(2026). HiLLM: Heterogeneous CPU–GPU Inference for Interactive MoE LLM Serving on Local Systems. 2026 59th IEEE/ACM International Symposium on Microarchitecture (MICRO).

(2025). CAPA: Convertibility-Aware Mixed-Precision Accelerator for Asymmetric GEMM in Weight-Only Quantized LLMs. 2026 63rd ACM/IEEE Design Automation Conference (DAC).

(2025). Survey and Evaluation of Converging Architecture in LLMs Based on Footsteps of Operations. IEEE Open Journal of the Computer Society (OJCS).

PDF Cite DOI

(2025). DBC: Drift-aware Binary Code for Drift-tolerant Deep Neural Networks. 2025 62th ACM/IEEE Design Automation Conference (DAC).

PDF Cite DOI

(2025). Bit-slice Architecture for DNN Acceleration with Slice-level Sparsity Enhancement and Exploitation. 2025 31st IEEE International Symposium on High-Performance Computer Architecture (HPCA).

PDF Cite DOI

(2024). ViT-slice: End-to-end Vision Transformer Accelerator with Bit-slice Algorithm. 2024 61th ACM/IEEE Design Automation Conference (DAC).

PDF Cite DOI

(2023). RQ-DNN: Reliable Quantization for Fault-tolerant Deep Neural Networks. 2023 60th ACM/IEEE Design Automation Conference (DAC) Late Breaking Results.

PDF Cite DOI

(2023). PIE-DRAM: Postponing IECC to Enhance DRAM performance with access table. 2023 60th ACM/IEEE Design Automation Conference (DAC) Late Breaking Results.

PDF Cite DOI

(2022). Bipolar vector classifier for fault-tolerant deep neural networks. 2022 59th ACM/IEEE Design Automation Conference (DAC).

PDF Cite DOI