Graph
Statistics
Wiki
Dark
wiki
/
concepts
/ 多模态模型
Type:
concept
Confidence:
0.70
Created:
2026-04-20
Updated:
2026-04-20
Tags:
multimodal
AI
AI工程
多模态模型
概述
能够处理和生成多种模态(文本、图像、音频等)数据的 AI 模型。
关键内容
GPT-4
:支持文本和图像输入,实现跨模态理解和推理。
CLIP
:建立文本和视觉的统一表示空间,实现图文对齐。
DALL-E
:从文本生成图像,实现跨模态生成。
来源
ai_papers_timeline.md
— 2021、2023 年时间线条目
相关
GPT 系列
— implemented_by
CLIP
— implemented_by
DALL-E
— implemented_by
← 多模态检索
多步回报 →