青青草a国产免费观看|91麻豆精品国产福利|国产av五无码一级毛片|亚洲爆乳精品无码一区二区|久久亚洲AV成人无码国产|91无码人妻一区二区三区|色婷婷av一区二区三区性色|国产制服91一区二区三区制服,女人书籍排行榜,盗墓笔记小说txt下载,玄幻小说排行榜完本

position: EnglishChannel  > AI ripples> Chinese AI Model Emu3 Handles Text, Image, Video Seamlessly

Chinese AI Model Emu3 Handles Text, Image, Video Seamlessly

Source: Science and Technology Daily | 2024-12-17 15:44:35 | Author: Gong Qian

On October 21, the Beijing Academy of Artificial Intelligence (BAAI), a Chinese non-profit organization engaged in AI R&D, released Emu3, a multimodal AI model that seamlessly integrates text, image, and video modalities into a single, unified framework.

The BAAI research team said Emu3 is expected to be used in scenario applications such as robot brains, autonomous driving, multimodal dialogue and inference.

Emu3, based solely on next-token prediction, proves that next-token prediction can be a powerful paradigm for multimodal models.

The existing multimodal AI models are mostly designed for specific tasks. Each has its corresponding architecture and methods. For instance, in the field of video generation, many developers use the diffusion in time (DiT) architecture, as referenced by Sora. Other models such as Stable Diffusion are used for text-to-image synthesis, Sora for text-to-video conversion, and GPT-4V for image-to-text generation.

In contrast to these models, which have a combination of isolated skills rather than an inherently unified ability, Emu3, eliminates the need for diffusion or compositional approaches. By tokenizing images, text, and videos into a discrete space, BAAI has developed a single transformer from scratch.

Emu3 outperforms several well-established task-specific models in both generation and perception tasks, surpassing flagship models such as SDXL and LLaVA.

In September, BAAI open-sourced the key technologies and models of Emu3 including the chat model and generation model after supervised fine-tuning.

Emu3 has been receiving rave reviews from overseas developers. "For researchers, a new opportunity has emerged to explore multimodality through a unified architecture, eliminating the need to combine complex diffusion models with large language models. This approach is akin to the transformative impact of transformers in vision-related tasks," AI consultant Muhammad Umair said on social media platform Meta.

While next-token prediction is considered a promising path towards artificial general intelligence, it struggled to excel in multimodal tasks, which were dominated by diffusion models such as Stable Diffusion and compositional approaches like CLIP combined with large language models.

Raphael Mansuy, co-founder of QuantaLogic, an AI agent platform, thinks that Em3 has significant implications for Al development. Mansuy wrote on X that Em3's success suggests several key insights: Next-token prediction as a viable path to general multimodal Al; potential for simplified and more scalable model architectures; challenge to the dominance of diffusion and compositional approaches.

Editor:GONG Qian

Top News

  • It is necessary to promote the opening up and sharing of scientific research infrastructure, make good use of multilateral mechanisms, and establish and improve international open sharing platforms, Chen Jiachang, China’s vice minister of science and technology, said at the Open Science International Forum, part of the 2025 Zhongguancun Forum Annual Conference, on March 28.

China Unveils Landmark AI-assisted Academic Monograph at London Book Fair

China's first AI-assisted academic monograph, AI for Rock Dynamics, was officially released at the London Book Fair on Mar 12, 2025. This groundbreaking work, led by Academician He Manchao, President of the Chinese Society for Rock Mechanics and Engineering (CSRME), and featuring contributions from 25 young scholars, demonstrates the growing integration of AI into academic research and publishing.

Chinese Museums Revive Millennia-old Civilizations with Digital Tech

According to data released by the National Cultural Heritage Administration (NCHA), museums across China received approximately 72.65 million visits from January 29 to February 4 this year, the first seven days of the Spring Festival holidays, with daily attendance increasing by 12.84 percent compared to the previous year.

抱歉,您使用的瀏覽器版本過低或開啟了瀏覽器兼容模式,這會影響您正常瀏覽本網(wǎng)頁

您可以進行以下操作:

1.將瀏覽器切換回極速模式

2.點擊下面圖標升級或更換您的瀏覽器

3.暫不升級,繼續(xù)瀏覽

繼續(xù)瀏覽
巴彦淖尔市| 建瓯市| 蒲江县| 鄂州市| 景德镇市| 兴山县| 顺平县| 永城市| 平凉市| 龙游县| 临洮县| 集安市| 青川县| 英山县| 曲靖市| 开阳县| 时尚| 东阳市| 开封县| 阿克陶县| 古浪县| 中阳县| 濉溪县| 资溪县| 襄垣县| 迁西县| 舟山市| 营山县| 乌海市| 嘉祥县| 玉环县| 龙门县| 嫩江县| 年辖:市辖区| 义乌市| 永和县| 邯郸市| 江华| 南投市| 潢川县| 隆尧县|