|
2024/10/30

Lei Li
Carnegie Mellon University
Web:https://www.cs.cmu.edu/~leili/
|
Lei Li is an Assistant Professor in Language Technologies Institute at Carnegie Mellon University. His research focuses on machine translation, trustworthy LLMs, and AI drug discovery. He received Ph.D. from CMU School of Computer Science in 2011. He is a recipient of the ACL 2021 Best Paper Award, the CCF Young Elite Award in 2019, the CCF Distinguished Speaker in 2017, the Wu Wen-tsün AI prize in 2017, and the 2012 ACM SIGKDD dissertation award (runner-up), and is recognized as Notable Area Chair of the ICLR 2023. Previously, he was an associate professor (tenured) at UC Santa Barbara. Before that, he was the Founding Director of ByteDance AI Lab, a principal scientist at Baidu, and a postdoc researcher at UC Berkeley. He led and developed ByteDance’s machine translation system VolcTrans and AI writing system Xiaomingbot, and many of his algorithms have been deployed in products (Toutiao, Douyin, Tiktok, Lark), serving over one billion users.
|
Can large language models (LLMs) fairly evaluate and refine model generation performance? Does a large language model reliably know specific knowledge, or does it answer by luck? In this talk, we will present scientific methods for evaluating large language models for knowledge-intense and language-generation tasks. We observe self-bias when using LLMs as evaluators—an LLM favors its own output. We will further discuss post-training methods to refine and align LLMs better with human judgment and valuation.
|
|
2024/12/4

Hongxia Yang
The Hong Kong Polytechnic University (PolyU)
|
Hongxia Yang, with over 15 years of experience as an AI scientist, specializes in large-scale machine learning, data mining, and deep learning. Throughout her career, she has developed 10 significant algorithmic systems, improving the operations of various enterprises. Her research includes pre-trained models, big data analytics, and the practical deployment of large language model(LLM) systems in real settings. Prof. Yang has published more than 100 top-tier papers, amassed around 10K citations with an H-index of 46, and holds over 50 patents. She has received several awards, including the 2019 SAIL Award at the World Artificial Intelligence Conference and the 2020 National Science and Technology Progress Award, China’s top tech accolade. Named one of Forbes China’s Top 50 Women in Tech in 2022 and AI 2000 Most Influential Scholar Award in 2023-2024, Prof. Yang has held prominent roles at ByteDance US, Alibaba Group, Yahoo! Inc, and IBM T.J. Watson Research Center. She earned her PhD from Duke University and her B.S. from Nankai University.
|
Title: Collaboration and Evolution of Foundation and Specialized Models
Abstract: The prevailing computing resource monopoly significantly restricts AI development, confining participation in the pretraining stages of Large Language Models (LLMs) to a few researchers. We are currently developing an efficient continual pretraining infrastructure designed to produce high-quality small language models (SLM) and multimodal SLM, with a particular emphasis on enhancing reasoning capabilities. We also introduces a novel system that integrates hundreds of domain-specific models to construct a foundational model for Artificial General Intelligence (AGI) with minimal computational demand. By employing smaller, efficient models, leveraging top-ranked models across diverse domains through a robust ranking algorithm, and continuously optimizing the evolving foundation model, this approach seeks to democratize AI development. It shifts from the traditional 'model over data' or centralized LLM method to a 'model over models' or decentralized LLM strategy, aiming to reduce reliance on extensive computational resources and promote broader innovation and inclusivity in AI.
|