|
Yuk-Kit Chan
Hello, I am a master's student in Computer Science and Engineering at the University of California, San Diego.
My research focuses on large language models, agent safety and factuality evaluation.
I earned my Bachelor of Engineering in Computer Science and Engineering from The Chinese University of Hong Kong,
where I completed the ELITE stream, and graduated with First Class Honours and a minor in Mathematics. I was very fortunate to have been advised by Professor
Wenxuan Wang.
Email  / 
Github / 
Google Scholar / 
CV / 
linkedin
|
|
University of California, San Diego, La Jolla, CA Sep 2026 - Jun 2028 (Expected)
Master of Science, Department of Computer Science and Engineering
Computer Science
|
The Chinese University of Hong Kong, Hong Kong Aug 2022 - Jul 2026
Bachelor of Engineering, Department of Computer Science and Engineering
Artificial Intelligence – Systems & Technologies, First Class Honours, ELITE Stream, Minor in Mathematics
|
|
|
Position: Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility.
Jialun Cao, Yuk-Kit Chan*, Zixuan Ling*, Wenxuan Wang, Shuqing Li,
Mingwei Liu, Chaozheng Wang, Boxi Yu, Pinjia He, Shuai Wang,
Zibin Zheng, Michael R. Lyu, Shing-Chi Cheung
ICML, 2026
arXiv |
ICML |
We conducted a decade-scale survey of 672 code benchmarks, found a growing awareness gap in rigorous benchmark practice, and respond by proposing HOW2BENCH
|
|
|
Identifying the Achilles' Heel: An Iterative Method for Dynamically Uncovering Factual Errors in Large Language Models
Wenxuan Wang, Yuk-Kit Chan*, Zixuan Ling*, SHI Juluan*,
Youliang Yuan, Jen-tse Huang, Wenxiang Jiao, Zhaopeng Tu, Michael R. Lyu
Findings of ACL, 2026
arXiv |
ACL |
Github
We introduce HalluHunter, a fully automated knowledge-graph framework that generates single- and multi-hop factuality questions and adaptively targets LLM weaknesses
|
|
|
Learning to Ask: When LLM Agents Meet Unclear Instruction.
Wenxuan Wang, Shi Juluan*, Zixuan Ling*, Yuk-Kit Chan*, Chaozheng Wang,
Cheryl Lee, Youliang Yuan, Jen-tse Huang, Wenxiang Jiao, and Michael R. Lyu
EMNLP, 2025
arXiv |
EMNLP |
Github |
We study tool-use under imperfect user instructions by analyzing real queries and introducing Noisy ToolBench,
showing LLMs often hallucinate missed arguments;
we then propose Ask-when-Needed (AwN) and an automated ToolEvaluator that improve tool-use accuracy over existing methods.
|
Last updated Sep 2026.
Website template from Jon
Barron.
|
|