图像理解
约 639 字大约 2 分钟
布欧-Lewyon
2026-05-15
首页 › Spring AI › 多模态与 Embedding
多模态能力让 AI 不仅能处理文本,还能理解图像、音频等内容。
图像理解基础
GPT-4o 和 GPT-4o-mini 支持图像输入,可以通过 Spring AI 发送图片并让 AI 描述或分析。
spring:
ai:
openai:
api-key: ${OPENAI_API_KEY}
chat:
options:
model: gpt-4o-mini # 需要支持图像的多模态模型发送图片 URL
@RestController
@RequestMapping("/api/vision")
public class VisionController {
private final ChatClient chatClient;
@GetMapping("/describe")
public String describeImage(@RequestParam String imageUrl) {
return chatClient.call(new Prompt(List.of(
new UserMessage(
"请用中文详细描述这张图片的内容",
List.of(new Media(MediaType.IMAGE, imageUrl))
)
))).getResult().getOutput().getContent();
}
}发送 Base64 图片
@PostMapping("/analyze")
public String analyzeImage(@RequestParam("file") MultipartFile file) throws IOException {
// 将上传的图片转为 Base64
String base64 = Base64.getEncoder().encodeToString(file.getBytes());
String dataUri = "data:%s;base64,%s".formatted(file.getContentType(), base64);
return chatClient.call(new Prompt(List.of(
new SystemMessage("You are a visual analyst. Describe what you see in detail."),
new UserMessage(
"分析这张图片",
List.of(new Media(MediaType.IMAGE, dataUri))
)
))).getResult().getOutput().getContent();
}图像分析应用
OCR 识别
@PostMapping("/ocr")
public String extractText(@RequestParam("file") MultipartFile image) throws IOException {
String base64 = Base64.getEncoder().encodeToString(image.getBytes());
String dataUri = "data:image/png;base64," + base64;
return chatClient.call(new Prompt(List.of(
new SystemMessage("你是 OCR 识别专家。提取图片中的所有文字,保持原格式输出。"),
new UserMessage(
"提取这张图片中的所有文字,包括标题、正文和标注。",
List.of(new Media(MediaType.IMAGE, dataUri))
)
))).getResult().getOutput().getContent();
}商品图分析
@PostMapping("/product/analyze")
public ProductAnalysis analyzeProductImage(@RequestParam("file") MultipartFile image) {
String base64 = Base64.getEncoder().encodeToString(image.getBytes());
String dataUri = "data:image/jpeg;base64," + base64;
BeanOutputConverter<ProductAnalysis> converter =
new BeanOutputConverter<>(ProductAnalysis.class);
String json = chatClient.call(new Prompt(List.of(
new SystemMessage("你是电商商品分析师。分析商品图片并提取结构化信息。"),
new UserMessage(
"分析这张商品图片,提取以下信息:\n" +
"1. 商品类别\n2. 主要颜色\n3. 品牌标识\n4. 预估价格区间\n5. 适用场景",
List.of(new Media(MediaType.IMAGE, dataUri))
)
), ChatOptionsBuilder.builder().withTemperature(0.3).build()));
return converter.convert(json);
}
public record ProductAnalysis(
String category,
String mainColor,
String brand,
String estimatedPriceRange,
String suitableScenario
) {}多图分析
@PostMapping("/compare")
public String compareImages(@RequestParam("files") List<MultipartFile> files) throws IOException {
List<Media> medias = new ArrayList<>();
for (MultipartFile file : files) {
String base64 = Base64.getEncoder().encodeToString(file.getBytes());
medias.add(new Media(MediaType.IMAGE,
"data:image/jpeg;base64," + base64));
}
return chatClient.call(new Prompt(List.of(
new UserMessage(
"比较这些图片,告诉我它们的异同点。",
medias
)
))).getResult().getOutput().getContent();
}小结
- GPT-4o / GPT-4o-mini 等支持图像理解,通过
UserMessage+Media传递。 - 支持 URL 和 Base64 两种图片传递方式。
MultipartFile上传配合 Base64 编码处理用户上传图片。- 结合
BeanOutputConverter可提取结构化图像分析结果。 - 多张图片可同时发送,用于对比分析。
上一节:数据库与 API Tool 实战 下一节:Embedding 向量化
