流式调用与参数调节
约 584 字大约 2 分钟
布欧-Lewyon
2026-05-15
首页 › Spring AI › ChatClient 与 Prompt 工程
流式调用
流式调用将 AI 的回复以流的方式逐步推送,用户体验更好(类似 ChatGPT 打字效果)。
StreamingChatClient
@Service
public class StreamingService {
private final StreamingChatClient streamingChatClient;
public StreamingService(StreamingChatClient streamingChatClient) {
this.streamingChatClient = streamingChatClient;
}
public Flux<String> stream(String message) {
return streamingChatClient.stream(message);
}
}SseEmitter(传统 MVC)
@RestController
public class SseController {
private final StreamingChatClient streamingChatClient;
@GetMapping("/chat/stream")
public SseEmitter streamChat(@RequestParam String message) {
// 创建 SseEmitter,超时时间 5 分钟
SseEmitter emitter = new SseEmitter(300_000L);
streamingChatClient.stream(message)
.subscribe(
content -> {
try {
emitter.send(SseEmitter.event()
.data(content)
.name("message"));
} catch (IOException e) {
emitter.completeWithError(e);
}
},
emitter::completeWithError,
emitter::complete
);
return emitter;
}
}WebFlux(响应式)
@RestController
public class WebFluxController {
private final StreamingChatClient streamingChatClient;
@GetMapping(value = "/chat/flux", produces = MediaType.TEXT_EVENT_STREAM_VALUE)
public Flux<ServerSentEvent<String>> streamFlux(@RequestParam String message) {
return streamingChatClient.stream(message)
.map(content -> ServerSentEvent.<String>builder()
.data(content)
.event("message")
.build())
.doOnError(e -> log.error("Stream error", e));
}
}前端集成(EventSource)
<!DOCTYPE html>
<html>
<body>
<div id="output"></div>
<script>
const source = new EventSource('/chat/flux?message=写一首诗');
source.addEventListener('message', (event) => {
document.getElementById('output').innerText += event.data;
});
source.onerror = () => source.close();
</script>
</body>
</html>参数调节(ChatOptions)
每个 AI 调用都可以指定参数控制模型行为。
全局参数
spring:
ai:
ollama:
chat:
options:
model: llama3.2:1b
temperature: 0.7
maxTokens: 500
topP: 0.9
frequencyPenalty: 0.0
presencePenalty: 0.0按请求参数
// 方式一:通过 Prompt
ChatResponse response = chatClient.call(
new Prompt(
"Generate a creative story",
ChatOptionsBuilder.builder()
.withTemperature(0.9) // 创作模式
.withMaxTokens(1000)
.build()
)
);
// 方式二:通过 Builder 链式调用
String response = chatClient.call(
new Prompt("Explain Spring AI in simple terms",
OllamaOptions.builder()
.withModel("llama3.2:3b") // 临时换模型
.withTemperature(0.3) // 更精确
.build()
)
);支持参数
ChatOptions options = ChatOptionsBuilder.builder()
.withTemperature(0.7) // 0.0 ~ 2.0,随机性
.withMaxTokens(500) // 最大生成 Token 数
.withTopP(0.9) // 0.0 ~ 1.0,核采样
.withFrequencyPenalty(0.0) // -2.0 ~ 2.0,频率惩罚
.withPresencePenalty(0.0) // -2.0 ~ 2.0,存在惩罚
.withStop(List.of("\n", ".")) // 停止序列
.build();参数调节策略
// 精确问答(低温度)
ChatOptions forFacts = ChatOptionsBuilder.builder()
.withTemperature(0.1)
.withTopP(0.1)
.build();
// 创意写作(高温度)
ChatOptions forCreative = ChatOptionsBuilder.builder()
.withTemperature(0.9)
.withTopP(0.9)
.build();
// 代码生成(中等温度)
ChatOptions forCode = ChatOptionsBuilder.builder()
.withTemperature(0.3)
.withMaxTokens(2000)
.build();小结
- 流式调用用
StreamingChatClient+Flux<String>,结合 SseEmitter 或 WebFlux 推送给前端。 - EventSource 在前端逐段接收 AI 响应,实现打字效果。
ChatOptions控制 Temperature / MaxTokens / TopP 等参数。- 低 Temperature(0.1)适合精确问答,高 Temperature(0.9)适合创意生成。
- 全局配置在 YAML,按需覆盖通过
Prompt的第二个参数传入。
上一节:Message 类型与多轮对话 下一节:结构化输出基础
