实体抽取实战
约 587 字大约 2 分钟
布欧-Lewyon
2026-05-15
实体抽取是 AI 最实用的能力之一——从非结构化文本中提取结构化信息。
简历信息抽取
public record Resume(
String name,
String email,
String phone,
List<String> skills,
List<Experience> experiences,
Education education
) {}
public record Experience(
String company,
String position,
String startDate,
String endDate,
List<String> responsibilities
) {}
public record Education(
String school,
String degree,
String major,
String graduationYear
) {}@Service
public class ResumeExtractionService {
private final ChatClient chatClient;
public ResumeExtractionService(ChatClient.Builder builder) {
this.chatClient = builder.build();
}
public Resume extractResume(String resumeText) {
BeanOutputConverter<Resume> converter = new BeanOutputConverter<>(Resume.class);
String response = chatClient.call(new Prompt("""
You are a resume parser. Extract structured information from the following resume.
Rules:
- Parse ALL experiences mentioned
- Normalize dates to YYYY-MM format
- If a field is not found, use null
- Don't miss any skill
Resume:
---
%s
---
%s
""".formatted(resumeText, converter.getFormatInstruction())));
return converter.convert(response);
}
}邮件意图分类
public record EmailIntent(
String category, // complaint / inquiry / request / feedback / spam
String summary, // 一句话摘要
String urgency, // low / medium / high
List<String> actionItems,
String replySuggestion // 建议回复
) {}@PostMapping("/email/classify")
public EmailIntent classifyEmail(@RequestBody String emailContent) {
BeanOutputConverter<EmailIntent> converter =
new BeanOutputConverter<>(EmailIntent.class);
String json = chatClient.call(new Prompt("""
Classify the following email and extract intent.
Email:
---
%s
---
%s
""".formatted(emailContent, converter.getFormatInstruction())));
return converter.convert(json);
}批量抽取
public record BatchResult<T>(List<T> items, int totalCount, List<String> errors) {}
public List<Person> batchExtract(List<String> texts) {
BeanOutputConverter<List<Person>> converter =
new BeanOutputConverter<>(new TypeReference<>() {});
// 分批处理
List<Person> results = new ArrayList<>();
for (int i = 0; i < texts.size(); i += 10) {
List<String> batch = texts.subList(i, Math.min(i + 10, texts.size()));
String joined = String.join("\n---\n", batch);
String json = chatClient.call(new Prompt("""
Extract person information from each entry separated by ---.
Entries:
%s
Return a JSON array of persons.
%s
""".formatted(joined, converter.getFormatInstruction())));
results.addAll(converter.convert(json));
}
return results;
}带验证的抽取
@Component
public class ValidatedExtractionService {
public <T> T extractAndValidate(String text, Class<T> type, Validator<T> validator) {
BeanOutputConverter<T> converter = new BeanOutputConverter<>(type);
int maxRetries = 3;
for (int i = 0; i < maxRetries; i++) {
String json = chatClient.call(new Prompt("""
Text: %s
%s
Important: Ensure ALL fields are correctly populated.
""".formatted(text, converter.getFormatInstruction())));
try {
T result = converter.convert(json);
if (validator.isValid(result)) {
return result;
}
log.warn("Validation failed, retrying ({}/{})", i + 1, maxRetries);
} catch (Exception e) {
log.error("Parse failed, retrying ({}/{})", i + 1, maxRetries, e);
}
}
throw new ExtractionException("Failed to extract valid entity after " + maxRetries + " retries");
}
}
// 校验器
interface Validator<T> {
boolean isValid(T entity);
}小结
- 实体抽取的核心模式:定义 Record →
BeanOutputConverter→ Prompt 指导 → 解析 JSON。 - 简历抽取、邮件分类是经典场景,支持嵌套对象和列表。
- 批量抽取时注意上下文窗口限制,分批处理。
- 生产环境需要加验证和重试机制,确保输出质量。
TypeReference处理List<Person>等泛型类型。
上一节:结构化输出基础 下一节:Tool Calling 基础
