Yuk-Kit Chan

Hello, I am a master's student in Computer Science and Engineering at the University of California, San Diego. My research focuses on large language models, agent safety and factuality evaluation.

I earned my Bachelor of Engineering in Computer Science and Engineering from The Chinese University of Hong Kong, where I completed the ELITE stream, and graduated with First Class Honours and a minor in Mathematics. I was very fortunate to have been advised by Professor Wenxuan Wang.

Email  /  Github /  Google Scholar /  CV /  linkedin

profile photo
Education
University of California, San Diego, La Jolla, CA Sep 2026 - Jun 2028 (Expected)
Master of Science, Department of Computer Science and Engineering
Computer Science
The Chinese University of Hong Kong, Hong Kong Aug 2022 - Jul 2026
Bachelor of Engineering, Department of Computer Science and Engineering
Artificial Intelligence – Systems & Technologies, First Class Honours, ELITE Stream, Minor in Mathematics
Research Experience
Position: Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility.
Jialun Cao, Yuk-Kit Chan*, Zixuan Ling*, Wenxuan Wang, Shuqing Li,
Mingwei Liu, Chaozheng Wang, Boxi Yu, Pinjia He, Shuai Wang,
Zibin Zheng, Michael R. Lyu, Shing-Chi Cheung
ICML, 2026
arXiv | ICML |

We conducted a decade-scale survey of 672 code benchmarks, found a growing awareness gap in rigorous benchmark practice, and respond by proposing HOW2BENCH

Identifying the Achilles' Heel: An Iterative Method for Dynamically Uncovering Factual Errors in Large Language Models
Wenxuan Wang, Yuk-Kit Chan*, Zixuan Ling*, SHI Juluan*,
Youliang Yuan, Jen-tse Huang, Wenxiang Jiao, Zhaopeng Tu, Michael R. Lyu
Findings of ACL, 2026
arXiv | ACL | Github

We introduce HalluHunter, a fully automated knowledge-graph framework that generates single- and multi-hop factuality questions and adaptively targets LLM weaknesses

Learning to Ask: When LLM Agents Meet Unclear Instruction.
Wenxuan Wang, Shi Juluan*, Zixuan Ling*, Yuk-Kit Chan*, Chaozheng Wang,
Cheryl Lee, Youliang Yuan, Jen-tse Huang, Wenxiang Jiao, and Michael R. Lyu
EMNLP, 2025
arXiv | EMNLP | Github |

We study tool-use under imperfect user instructions by analyzing real queries and introducing Noisy ToolBench, showing LLMs often hallucinate missed arguments; we then propose Ask-when-Needed (AwN) and an automated ToolEvaluator that improve tool-use accuracy over existing methods.





Last updated Sep 2026.
Website template from Jon Barron.