Software Development, AI
Event-driven architecture: 4 patterns and broker vs mediator
TL;DR: Event-driven architecture lets services talk by announcing what happened, so each one can scale, fail and ship on its own. Keep events lean and versioned, pick broker topology for simple independent steps and mediator for workflows that must retry or roll back, and plan for duplicates and short delays. Skip it when plain request-response calls already do the job.
Event-driven architecture (EDA) is a way to build software where services talk by saying what happened, not by calling each other. One service says "order placed". Every service that cares picks it up and gets on with its job. The one that raised the event doesn't know who's listening. It doesn't need to.
That's the whole trick. Each part can scale, fail and ship on its own. But it has real costs, so it's not the right call for every build.
So here's how it works, the four event patterns worth knowing, and how to pick between broker and mediator.
How event-driven architecture works
Every event-driven system has three jobs. A producer raises an event. A channel carries it. One or more consumers react. Microsoft's Azure guide to event-driven architecture sets it out the same way, with the channel often run by an event broker.
Compare that with how most software gets built. Gregor Hohpe's Enterprise Integration Patterns paper on events calls it command-and-control: one method calls another and tells it what to do. The real world doesn't work like that. An order comes in, and a few systems react to it in their own time.
The channel you pick matters more than most teams think. A simple publish-subscribe queue hands a message to whoever is listening, then forgets it. A log-based platform keeps it. In Apache Kafka, events aren't deleted once they're read. You set how long to keep them. So a new consumer can join next month and replay history. Confluent's EDA overview makes the same point: a durable event store gives you an audit trail, and replaying events is how you get back to a known state after a failure.
Two things catch people out early.
Order is the first. Kafka only promises order inside a partition. Events with the same key land in the same partition, so your key decides what "in order" means. Key on customer ID and each customer's events stay in order. Events across two customers don't.
Duplicates are the second. Plan for them. Azure warns that running many copies of a consumer causes trouble if your processing isn't idempotent. In plain words: handling the same event twice must do no harm. Build every consumer that way from day one.
Broker vs mediator topology
Azure names two ways to run the flow of events.
In a broker topology, services broadcast events to the whole system. Others act on them or ignore them. Nothing sits in the middle. It's the loosest setup you'll get, and Azure says it suits a simple flow. The catch? There's no built-in way to restart or replay a business process that failed halfway, and it can leave your data out of step.
In a mediator topology, an event mediator runs the flow. It keeps state, handles errors and can restart a process. It sends commands to the right channel for each step. You give up some independence. You get one place to see and fix the whole workflow.
Here's the simple test. Independent steps that don't care about each other? Broker. A multi-step process that must pause, retry or roll back cleanly? Mediator.
This is the same call as choreography vs orchestration, and we've gone deep on it in orchestration vs choreography: 6 criteria that decide it. If the mediator is an AI workflow, durable execution matters a lot. AI Orchestrators' guide to six AI orchestration patterns covers why.
Four event patterns worth knowing
The broker gets all the attention. What you put inside the event matters more. Martin Richards' guide to event design patterns puts it well: every event you publish is a contract with your consumers. He lays out four patterns.
- Event notification. A small message, often just an ID and a type. The consumer calls back for the detail. Low coupling, but lots of follow-up traffic.
- Event-carried state transfer. The event carries the data consumers need, so they don't call back. Less chatter. But now every consumer depends on that payload, so you can't change it on a whim.
- Change data capture (CDC). Database row changes get streamed out as events. It's a handy way to pull events out of an old system without rewriting it.
- Async commands. Not really an event. A command tells one service to do something. An event says something already happened. Mix the two up and the design gets confusing fast.
AWS's EDA overview draws the same line between the first two. An event can carry the state (the item, its price, the delivery address) or it can just be an identifier.
Watch for fat events. Richards calls them out as a common anti-pattern: events stuffed with data the sending service doesn't own. They feel handy at first. Later, every consumer is tied to data it should never have seen, and you can't untangle it.
Most real systems mix these patterns across different parts of the business. That's fine. Messy, unversioned schemas are the thing to fear, not variety.
Where event-driven architecture earns its keep
A few jobs come up again and again.
- Order fan-out. "Order placed" kicks off fulfilment, billing, stock and the customer email at once. AWS points to fan-out as a core use: the router pushes one event to many systems, each working in parallel, with no custom code per consumer.
- Device data. Thousands of sensors sending readings nonstop. A streaming platform soaks that up far better than an API that answers one request at a time.
- Moving data between regions and accounts. AWS lists this too, using an event router so teams can build and deploy on their own.
- Streaming ETL. RisingWave's streaming ETL page pitches continuous pipelines as the fix for batch jobs, with data ready in under a second. It names real-time feature serving and fraud detection as the reasons. Worth noting it's a vendor making the case. Our post on ETL vs ELT covers the batch side of that choice.
- AI features. An event can trigger a model call, or feed fresh data to a live model. AI-Led's LLM integration guide covers wiring AI into a product. Calling an AI API from an event consumer brings its own traps, like rate limits, retries and spend, and AI-Led's AI API integration guide walks through them.
The real benefits and the real costs
The upside is easy to sell. Each consumer scales to its own load without dragging the rest along. A durable log gives you replay and an audit trail for close to free.
Now here's the important bit. Data doesn't move instantly. For a short window, different parts of the system will disagree about what's true. Azure is blunt about it: if you can't live with that window, event-driven architecture works against you. Debugging gets harder too, because one failure can start three services away. And you've got more to run: brokers, schema checks and tracing.
Azure's answer to failed events is a pattern worth copying:
- Send a failed event to an error handler, and keep processing the rest.
- If the handler fixes it, resubmit the event. Expect it to land out of order.
- If it can't be fixed, park it in a dead-letter queue for a person to check.
- When a business process spans several services, use a compensating transaction to undo the steps that already ran.
Schemas are the part teams skip. Producers and consumers ship on their own schedules, so you can't update them all at once. Azure's advice: pick a versioning plan early, and build consumers that cope with a version they don't know yet.
And sometimes the answer is no. Azure says to skip EDA when simple request-response calls already meet your needs, when every service must agree on the data at all times, or when your team has never run a distributed async system.
How to start without betting the platform
Don't rip out a working system to go all-in. Start small, prove it, then grow.
- Pick one bounded context. That's one part of the business with clear edges, like orders or billing.
- Write the event contracts first. Field names, types, what's required, and who owns changes.
- Version the schema from the very first event, not after the first break.
- Add tracing on day one. Put a correlation ID on every event so you can follow a failure back to where it started. Our guide to microservices architecture covers observability across services.
- Use CDC on your existing database to peel events out bit by bit, instead of rewriting it in one risky release.
- Plan replay before you need it. How will you reprocess events after a bug fix?
- Test in staging with traffic shaped like production. Replay bugs show up fast under load and hide in a quiet dev setup.
If you're weighing the wider scaling picture before you commit to a build, this scalable app development guide is a useful companion read on the ops side.
If your biggest pain is services waiting on each other at peak load, events are worth the extra moving parts. If the pain is slow queries or messy code, they're not. Fix that first.
*James*
Where Devwiz fits
Getting the topology, the contracts and the schema rules right early is the difference between a platform that scales cleanly and one that needs a rewrite in 18 months.
That's the work we do. We build AI platforms and programs as production software, with the data pipelines, event contracts and AI features that sit on top. For teams moving off an older system, our custom software development work covers the integration and step-by-step migration side.
Devwiz has shipped 200+ apps, including work for NSW Government (Justice and Corrective Services), Briometrix, Vivid and Huskee.
Scoping a platform that needs to react in real time, not poll for updates? Start with our services. Worth a chat.
Frequently asked questions
What is event-driven architecture?
It's a way to build software where services share events, facts about something that happened, instead of calling each other directly. A producer raises the event, a broker carries it, and any service that cares reacts. Nobody waits on anyone else.
How does event-driven architecture work?
A producer publishes an event, like "order placed", to a channel. Consumers subscribe to the events they care about and react in their own time. With a log-based platform like Kafka, events are kept for a set time, so a new consumer can replay history.
What is the difference between microservices and event-driven architecture?
Microservices is about splitting an app into services you can deploy on their own. Event-driven architecture is about how those services talk. You can run microservices over direct API calls or over events. Most mature systems use both.
What are the downsides of event-driven architecture?
Data isn't in sync straight away, so parts of the system can briefly disagree. Debugging across services is harder, and you have more to run: brokers, schema versioning and tracing. Microsoft's Azure guidance says to skip it if you need strong consistency or simple request-response already works.
Is Kafka an event-driven system?
Kafka is an event streaming platform, and a common base for event-driven systems. It stores events in partitions, keeps them after they're read, and guarantees order inside each partition. The architecture is the pattern of producers and consumers you build on top of it.
About James Killick
10+ years building digital products · 200+ apps shipped since 2015
James is a co-founder of Devwiz and an AI product specialist. Since 2015 he has helped ship 200+ apps for founders, businesses and government, including work for NSW Government, Briometrix and Huskee. He builds AI-first platforms and writes about turning a proven program into software. He also hosts the Up in the AI podcast.
More articles by James · James's personal site · LinkedIn · AI Orchestrators
Tags: Event-Driven, Software Architecture, Microservices, Data Pipelines


