| Prof. Sam KwongLingnan University (Hong Kong) IEEE Fellow Professor KWONG Sam Tak Wu is the Associate Vice-President (Strategic Research), J.K. Lee Chair Professor of Computational Intelligence and the Dean of the School of Graduate Studies of Lingnan University. Professor Kwong is a distinguished scholar in evolutionary computation, artificial intelligence (AI) solutions, and image/video processing, with a strong record of scientific innovations and real-world impacts. Professor Kwong is one of the most highly cited researchers by Clarivate in 2022, 2023, 2024 and 2025. He has also been actively engaged in knowledge transfer between academia and industry. He was elevated to IEEE Fellow in 2014 for his contributions to optimization techniques in cybernetics and video coding. He was the President of the IEEE Systems, Man, and Cybernetics Society (SMCS) in 2021-22. He is a fellow of US National Academy of Inventors (NAI), Canadian Academy of Engineering (CAE) and the Hong Kong Academy of Engineering (HKAE). Professor Kwong has a prolific publication record with over 500 journal articles, and 160 conference papers with an h-index of 99 based on Google Scholar. He is currently the associate editor of a number of leading IEEE transaction journals. |
| Prof. Yao ZhaoBeijing Jiaotong University, China IEEE Fellow Professor Zhao Yao is a recipient of the National High-Level Talent Program of the Ministry of Education, a Distinguished Young Scholar of the National Natural Science Foundation of China (NSFC), a recipient of the "Science and Technology Beijing" Top 100 Leading Talent Award, and a Fellow of the IEEE. He received his Ph.D. in Engineering from Beijing Jiaotong University in 1996, was exceptionally promoted to full professorship in 2001, and was appointed as a Ph.D. Supervisor in 2002. He is currently a Level-II Professor at Beijing Jiaotong University. Professor Zhao conducted postdoctoral research at Delft University of Technology (TU Delft), the Netherlands, from 2001 to 2002. He also held visiting positions at the University of Southern California, USA, and at EPFL (École Polytechnique Fédérale de Lausanne), Switzerland, in October 2015. He currently serves as the Director of the Beijing Key Laboratory of Intelligent Processing for Science-Fiction Audio and Video, and as the Director of the International Joint Research Laboratory for Cross-Modal Intelligent Innovation, jointly established by the Ministry of Education. His research focuses on a broad range of topics in digital media information processing and intelligent analysis, including artificial intelligence, advanced electronic information technology, software engineering, big data technology and engineering, generative artificial intelligence and security, as well as computer technology. Speech Title: Generalized and Explainable AI-Generated Image Detection Abstract: The rapid advancements in generative AI have significantly impacted digital forensics and security, enabling the creation of highly realistic synthetic images that closely resemble authentic visuals. While these innovations present transformative opportunities for creative industries, they also introduce substantial challenges in detecting AI-generated content and preventing its misuse in misinformation campaigns and fraudulent activities. In this talk, we will highlight our recent efforts to advance AI-generated image detection, focusing on improving detection accuracy and enhancing generalization across a range of generation methods. We will also explore the development of explainable AI-driven solutions, underscoring the urgent need for robust, scalable, and interpretative approaches to mitigate the growing threats posed by AI-generated content. |
| Prof. Siwei MaPeking University, China IEEE Fellow Dr. Siwei Ma isaBoya Distinguished Professor at Peking University, IEEE Fellow. His main research interests lie in video processing and coding. Heserves(served) as an Associate Editor for IEEE Transactions on Image Processing, IEEE Transactions on Circuits and Systems for Video Technology, and Journal of Visual Communication and Image Representation, as well as the Co-Chair of the Program Committee for IEEE VCIP 2017. Speech Title: Video coding, understanding and generation Abstract: Traditional video understanding has evolved from handcrafted feature extraction to Transformer-based multimodal large models, enabling the parsing and extraction of high-level semantics from raw pixels. Video generation has undergone two major developmental phases: Generative Adversarial Networks (GANs) and diffusion models, and is now advancing toward general world simulators with complete physical constraints. Starting from classic hybrid coding frameworks, video coding has sequentially spawned neural network-aided coding and end-to-end deep coding solutions, and is now entering a brand-new era of generative coding. The three technologies share a unified spatio-temporal semantic representation system and form a tight closed loop. Video understanding extracts high-level semantics such as objects, motions and scenes from frames; video generation reconstructs high-fidelity visual scenes conditioned on the extracted semantics; video coding preserves human perceptual quality and task-effective information for machine vision with minimal bit overhead. Centered on the internal correlations among video understanding, video generation and video coding, this report systematically sorts out recent research progress in this field, and focuses on cutting-edge technical routes including deep learning hybrid coding, end-to-end generative compression, and dedicated feature coding tailored for machine vision. |
| Prof. Cheng DengHohai University, China Prof. Deng secured over 30 research grants from national and provincial funding bodies. In the past five years, he published more than 200 papers in CCF-A journals and conferences, including ICML, NeurIPS, ICLR, CVPR, and ICCV. With over 16,000 citations on Google Scholar, he has been consecutively recognized as a Highly Cited Researcher in China and among the World's Top 2% Scientists. Hisresearch achievements have been honored with Second Prize of the National Natural Science Award and First Prizes of the Shaanxi Provincial Science and Technology Award. Speech Title:From Known to Unknown: Knowledge Evolution in the Open World Abstract:This talk will systematically elaborate on the cognitive evolution pathways of AI systems in open environments. It introduces a comprehensive three-stage cognitive closed-loop framework, i.e., Open-Set Recognition, Knowledge Discovery, and Continual Learning. Specifically, this talk details our research achievements, including the Dual-Stream Information Bottleneck architecture, Attribute Pool construction, and Memory Replay mechanisms. These achievements empower AI systems to recognize unknown entities, acquire novel knowledge, and mitigate catastrophic forgetting. We hope our achievements can provide a crucial theoretical framework and practical roadmap for achieving continual learning and advancing toward Artificial General Intelligence (AGI). |
| Prof. Weiming HuInstitute of Automation, Chinese Academy of Sciences (CAS),ChinaProfessor Hu Weiming is a Distinguished Core Research Scientist and Doctoral Supervisor at the National Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences (CAS). He serves as the Head of the Video Content Security Research Team and the Director of the Beijing Key Laboratory of Multimodal Superintelligence Security. Professor Hu is a recipient of the National Outstanding Youth Science Fund, a selected member of the "Ten Thousand Talents Program" for Science and Technology Innovation Leading Talents under the Organization Department of the CPC Central Committee, a selected member of the Young and Mid-aged Science and Technology Innovation Leading Talents under the Ministry of Science and Technology, a national-level selected candidate of the "Hundred Talents Program" under the Ministry of Human Resources and Social Security, a National Young and Mid-aged Expert with Outstanding Contributions, a recipient of the Special Government Allowance of the State Council, and the Chief Expert of the National Key Special Program on Information Security. He currently serves as an Associate Editor for IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI), ACM Transactions on Privacy and Machine Learning (PML), and IEEE Transactions on Cybernetics. His current research interests include visual content understanding and network multimedia sensitive content recognition. He has led over forty research projects, including Key Programs of the National Natural Science Foundation of China, National 863 Key Special Programs, and goal-oriented research projects. He has published over 400 papers in prestigious international journals such as PAMI and IJCV, as well as in leading domestic journals and major international conferences including NIPS and ICCV. He holds over 100 authorized invention patents. The sensitive multimedia recognition technologies developed by his team have been deployed in practical applications across more than 200 enterprises and institutions, delivering significant results in real-world operations and generating substantial economic and social benefits. As the first recipient, Professor Hu has been honored with the National Natural Science Award (Second Class), the Beijing Science and Technology Award (First Class, Technological Invention Category), the Beijing Invention Patent Award (First Class), and the Wu Wenjun Artificial Intelligence Science and Technology Award (First Class). Speech Title:Vision-Language Cross-Modal Pretraining and Matching Abstract:Mainstream pre-training methods for image-text retrieval typically adopt a dual-encoder architecture. However, improving the accuracy of dual-encoders remains challenging. This paper proposes a vision-language error modeling approach that injects fine-grained image-text association information into dual-encoders. Specifically, we leverage large language models to generate image captions with localized errors, constructing high-quality negative samples for training the dual-encoder. By detecting and rectifying textual errors based on visual information, the dual-encoder achieves enhanced extraction of fine-grained image and text features, while enabling text features to align with both global and local visual features. Furthermore, we propose a pre-training method based on contrastive local ranking distillation, which distills the ability to rank hard image-text pairs from more accurate yet computationally expensive joint pre-training models into the dual-encoder model, thereby improving its matching capability. Experimental results demonstrate that our proposed method outperforms state-of-the-art dual-encoder approaches on image-text retrieval benchmarks and significantly enhances the discriminative power for local textual semantics. |
```