Let's get started with a Microservice Architecture with Spring Cloud:
Building Graph-Based Workflows With Spring AI Alibaba
Last updated: October 10, 2026
1. Overview
When building AI agents, a single prompt to an LLM is usually not enough. Solving complex tasks requires multiple steps, where the output of one step decides what happens next. However, it’s difficult to manually implement and maintain this orchestration logic.
In this tutorial, we’ll solve this problem using Spring AI Alibaba. We’ll use it to model an agentic workflow as a graph of steps, which decides its own path at runtime.
2. Graph-Based Workflows 101
A graph-based workflow breaks a complex task into smaller units of work and connects them to form a graph. Instead of hardcoding the order of execution, we just describe the possible paths. The workflow follows the appropriate path based on the data it has.
Such workflows consist of three main concepts:
- State: the shared data that flows through the workflow.
- Nodes: the individual units of work in our graph. They typically read the state, perform business logic, and update the state at the end.
- Edges: the connections between the nodes. They control the order of execution by deciding which node to move to.
To see these concepts in action, we’ll build an excuse escalation workflow. First, an employee node invents an excuse for a given situation at work. Then, a manager node rates how believable the excuse is.
Next, an edge checks the manager’s rating. If the manager believes the excuse, the workflow ends. Otherwise, the flow loops back to the employee node, which refines the original excuse.
3. Configuring an LLM
Now, let’s set up our project. We’ll start by adding the necessary dependencies to our pom.xml file:
<dependency>
<groupId>org.springframework.ai</groupId>
<artifactId>spring-ai-starter-model-openai</artifactId>
<version>2.0.1</version>
</dependency>
<dependency>
<groupId>com.alibaba.cloud.ai</groupId>
<artifactId>spring-ai-alibaba-graph-core</artifactId>
<version>1.1.2.3</version>
</dependency>
Here, we first add Spring AI’s OpenAI starter dependency, which we’ll use to interact with an LLM. Next, we add the Spring AI Alibaba graph core dependency. This provides the classes we need to define our graph-based workflow.
To avoid version conflicts and compatibility issues between these dependencies, let’s also include the Spring AI Bill of Materials (BOM):
<dependencyManagement>
<dependencies>
<dependency>
<groupId>org.springframework.ai</groupId>
<artifactId>spring-ai-bom</artifactId>
<version>2.0.1</version>
<type>pom</type>
<scope>import</scope>
</dependency>
</dependencies>
</dependencyManagement>
Next, let’s configure our OpenAI API key and chat model in the application.yaml file:
spring:
ai:
openai:
api-key: ${OPENAI_API_KEY}
chat:
model: gpt-5.6-luna
Here, we specify OpenAI’s GPT-5.6 Luna model using the gpt-5.6-luna model ID. Alternatively, we can use a different chat model, as the specific AI model or provider is irrelevant for this demonstration.
With these two properties set, Spring AI automatically creates a ChatClient.Builder bean. We’ll use it to interact with the specified model in our nodes.
4. Defining State Keys
The state of our workflow is a collection of key-value pairs. Before our nodes can read or write them, we need to declare these keys along with a strategy. These strategies control how the graph updates a value.
Let’s define these keys in a KeyStrategyFactory bean:
@Bean
KeyStrategyFactory excuseKeyStrategyFactory() {
return () -> {
Map<String, KeyStrategy> strategies = new HashMap<>();
strategies.put("situation", new ReplaceStrategy());
strategies.put("excuse", new ReplaceStrategy());
strategies.put("managerReplies", new AppendStrategy());
strategies.put("believability", new ReplaceStrategy());
strategies.put("attempts", new ReplaceStrategy());
return strategies;
};
}
Here, the situation key holds the initial input of our workflow. The remaining keys represent the values our nodes will produce.
For most of them, we use the ReplaceStrategy, which overwrites the existing value with the new one. However, we use the AppendStrategy for the managerReplies key, which adds every new value to a list.
5. Implementing Nodes
With our state keys defined, let’s build the nodes of our graph. To create one, we implement the NodeAction interface and its apply() method. This method receives the current state in the form of an OverAllState instance. Then, after executing the business logic, we return a Map containing the updates we want to make to the state.
5.1. Inventing an Excuse
First, let’s create the node that plays the role of our employee:
@Component
class InventExcuseNode implements NodeAction {
private static final String PROMPT_TEMPLATE = """
You're an employee explaining why you missed something at work.
You're a bad liar, but a creative one.
You never take accountability or admit to lying.
Situation: {situation}
The excuse you already gave: {previousExcuse}
How your manager responded: {managerReplies}
Give a new excuse in at most two sentences. If you already gave one, do not
abandon it. Keep the original story and add a further complication.
Respond with only the excuse.
""";
private final ChatClient chatClient;
InventExcuseNode(ChatClient.Builder chatClientBuilder) {
this.chatClient = chatClientBuilder.build();
}
}
In our new class, we define the LLM instructions in a prompt template. Additionally, we inject the ChatClient.Builder bean that Spring AI auto-configures and use it to build a ChatClient instance. Next, let’s implement the apply() method:
@Override
public Map<String, Object> apply(OverAllState state) {
String situation = state.value("situation", String.class)
.orElseThrow(IllegalStateException::new);
String previousExcuse = state.value("excuse", "none yet");
List<String> managerReplies = state.value("managerReplies", List.of());
Integer attempts = state.value("attempts", 0);
String excuse = chatClient
.prompt()
.user(user -> user.text(PROMPT_TEMPLATE)
.param("situation", situation)
.param("previousExcuse", previousExcuse)
.param("managerReplies", managerReplies.isEmpty()
? "none yet"
: String.join("\n", managerReplies)))
.call()
.content();
return Map.of(
"excuse", excuse,
"attempts", attempts + 1
);
}
Here, we read the current state using the value() method. We define default values for these keys, except for the situation key. Instead, we treat it as a mandatory input to start the workflow.
Then, we fill these values into the prompt template, call the LLM, and get the excuse using the content() method. Finally, we update the state with the new excuse along with the incremented attempt count.
5.2. Reacting as the Manager
Next, we need a manager node that rates the employee’s excuse. Let’s create a new ManagerReactsNode class and implement the apply() method:
private static final String PROMPT_TEMPLATE = """
You are a tired engineering manager listening to an employee explain themselves.
Situation: {situation}
Their excuse: {excuse}
Reply in one sentence and rate how believable the excuse is, from 0.0 to 1.0.
""";
@Override
public Map<String, Object> apply(OverAllState state) {
String situation = state.value("situation", String.class)
.orElseThrow(IllegalStateException::new);
String excuse = state.value("excuse", String.class)
.orElseThrow(IllegalStateException::new);
ManagerReaction managerReaction = chatClient
.prompt()
.user(user -> user.text(PROMPT_TEMPLATE)
.param("situation", situation)
.param("excuse", excuse))
.call()
.entity(ManagerReaction.class);
return Map.of(
"managerReplies", managerReaction.reply(),
"believability", managerReaction.believability()
);
}
record ManagerReaction(
String reply,
Double believability
) {}
Similar to our previous node, we read the current state, populate the prompt template, and invoke the LLM. However, instead of calling the content() method, we call the entity() method to get a structured output. Finally, we return the reply and the believability score as state updates.
6. Creating a Conditional Edge
Now, we need a way to connect the nodes we’ve implemented. Additionally, we need to write logic to decide whether the workflow should end or the employee should try again. To do this, we’ll create a conditional edge by implementing the EdgeAction interface:
@Component
class ExcuseDispatcher implements EdgeAction {
private static final double BELIEVABILITY_THRESHOLD = 0.8;
private static final int MAX_ATTEMPTS = 3;
@Override
public String apply(OverAllState state) {
Double believability = state.value("believability", 0.0);
Integer attempts = state.value("attempts", 0);
if (isBelieved(believability) || attempts >= MAX_ATTEMPTS) {
return "stop";
}
return "escalate";
}
static boolean isBelieved(Double believability) {
return believability >= BELIEVABILITY_THRESHOLD;
}
}
Here, we read the believability and attempts keys from the state. If the manager’s rating reaches our configured threshold, we return the stop label. Otherwise, we return the escalate label to give the employee another chance.
In the first condition, we also check if the employee has exhausted their maximum attempts. This is important to avoid an infinite loop in our workflow.
7. Assembling the Graph
With all our building blocks in place, let’s wire them together and expose a CompiledGraph bean:
@Bean
CompiledGraph excuseGraph(
KeyStrategyFactory keyStrategyFactory,
InventExcuseNode inventExcuseNode,
ManagerReactsNode managerReactsNode,
ExcuseDispatcher excuseDispatcher
) throws GraphStateException {
StateGraph excuseGraph = new StateGraph(keyStrategyFactory)
.addNode("invent_excuse", AsyncNodeAction.node_async(inventExcuseNode))
.addNode("manager_reacts", AsyncNodeAction.node_async(managerReactsNode))
.addEdge(StateGraph.START, "invent_excuse")
.addEdge("invent_excuse", "manager_reacts")
.addConditionalEdges(
"manager_reacts",
AsyncEdgeAction.edge_async(excuseDispatcher),
Map.of(
"escalate", "invent_excuse",
"stop", StateGraph.END
)
);
return excuseGraph.compile();
}
First, we create a StateGraph instance and pass it our KeyStrategyFactory bean. This helps the graph handle updates for each state key. Next, we register our two nodes using the addNode() method, giving each of them a unique name.
Then, we define the flow using the addEdge() method. We connect the START constant to our invent_excuse node, marking it as the entry point. From there, we connect invent_excuse to the manager_reacts node.
After that, we attach our excuseDispatcher to the manager_reacts node using the addConditionalEdges() method. The map we pass here links the labels our dispatcher returns to their target nodes. The escalate label loops back to the invent_excuse node. Meanwhile, the stop label moves to the END constant, which finishes the workflow.
8. Testing the Workflow
Now that we’ve assembled our graph, let’s see how the workflow behaves under different situations.
8.1. Exposing a REST API
First, let’s expose an API endpoint that allows us to initiate the workflow:
@PostMapping("/excuse")
ResponseEntity<ExcuseResponse> generateExcuse(@RequestBody ExcuseRequest request) {
OverAllState finalState = excuseGraph
.invoke(
Map.of("situation", request.situation()),
RunnableConfig.builder()
.threadId(UUID.randomUUID().toString())
.build())
.orElseThrow();
ExcuseResponse response = new ExcuseResponse(
finalState.value("excuse", ""),
finalState.value("managerReplies", List.of()),
finalState.value("attempts", 0),
ExcuseDispatcher.isBelieved(finalState.value("believability", 0.0)));
return ResponseEntity.ok(response);
}
record ExcuseRequest(String situation) {}
record ExcuseResponse(
String finalExcuse,
List<String> managerReplies,
Integer attempts,
boolean believed
) {}
Here, we call the invoke() method on our injected CompiledGraph bean. We pass the situation from the request as the initial state. Additionally, we pass a RunnableConfig instance with a random threadId. This helps keep execution of concurrent requests separate from each other.
The invoke() method runs the graph from start to end and returns the final state. Then, we simply read values from this state and use them to prepare the API response.
8.2. Invoking the API Endpoint
Now, let’s use the HTTPie CLI to invoke our API endpoint:
http POST :8080/excuse situation="Employee was 30 seconds late to the daily standup"
Here, we enter a mundane situation. Let’s see what we get as a response:
{
"attempts": 1,
"believed": true,
"finalExcuse": "Sorry, I was only 30 seconds late because my calendar reminder fired late while my laptop was still reconnecting to the meeting room.",
"managerReplies": [
"Understood—30 seconds is minor, but please make sure you join promptly next time."
]
}
As we can see, the manager believes the very first excuse, so the workflow ends after a single attempt. Consequently, the managerReplies list contains just one reply, since the loop never ran again.
Next, let’s try a situation that’s much harder to explain and see the API response:
http POST :8080/excuse situation="The employee didn't show up to work for 3 days without any notice. Yet, the employee uploaded pictures of them partying on their public instagram account."
{
"attempts": 3,
"believed": false,
"finalExcuse": "The retreat’s emergency coordinator told us all communication had to go through their satellite system, but a solar flare knocked it offline and corrupted the evacuation logs. That also delayed the proof they promised to send confirming the phone confiscation, shuttle failure, and automated Instagram reposts.",
"managerReplies": [
"Frankly, this is difficult to believe, and disappearing for three days without finding any way to notify us is unacceptable.",
"That explanation is extraordinarily unlikely, and we need to discuss your three-day absence and the conflicting Instagram activity immediately.",
"That explanation strains credibility beyond reason; send the promised proof immediately, and we’ll discuss your three-day unreported absence."
]
}
Here, the manager doesn’t believe any of the excuses. Hence, the workflow loops until it reaches the maximum number of attempts.
9. Conclusion
In this article, we implemented a graph-based agentic workflow using Spring AI Alibaba.
We started by defining the state keys along with their strategies. Then, we implemented two LLM-powered nodes and a conditional edge that decides whether to loop or stop the workflow. Finally, we assembled everything into a graph and saw how the workflow adapts its path at runtime.
As always, all the code examples used in this article are available over on GitHub.
















