On this page
The ten protocols in one table
- REST
- GraphQL
- gRPC
- WebSockets
- Server-Sent Events
- Webhooks
- AMQP and MQTT
- MCP
- A2A
- WebRTC
The cheat sheet
How this appears in an interview
Frequently asked questions
Related reading
10 API Protocols Every Engineer Should Know


On This Page
The ten protocols in one table
- REST
- GraphQL
- gRPC
- WebSockets
- Server-Sent Events
- Webhooks
- AMQP and MQTT
- MCP
- A2A
- WebRTC
The cheat sheet
How this appears in an interview
Frequently asked questions
Related reading
Ten protocols cover almost every connection you will build: REST, GraphQL, gRPC, WebSockets, Server-Sent Events, webhooks, and the AMQP and MQTT message brokers. Three more have become unavoidable because AI agents run on them: MCP, A2A, and WebRTC.
Most engineers learn the first seven on the job and meet the last three only when an agent project lands on their desk. The two groups are not separate worlds. The agent protocols are built on the classic ones, so the failure you already know from a webhook or a stream is the failure you will meet again inside an agent system.
This article explains what each protocol does, when it is the right choice, and the specific way it breaks in production.
The ten protocols in one table
| Protocol | What it does | Use it for | How it breaks |
|---|---|---|---|
| REST | Resources and HTTP verbs, no session | Public CRUD APIs | A retried POST creates duplicates |
| GraphQL | The client asks for exact fields | Many screens, shifting data needs | One deep query overloads the database |
| gRPC | Typed contracts over HTTP/2 | Service-to-service calls | A missing deadline holds up the chain |
| WebSockets | One connection, both directions | Chat, multiplayer, live editing | Connections drop without an error |
| SSE | The server streams to the client | Streaming model output | A buffering proxy holds the stream |
| Webhooks | Another system calls your URL | Payments, job completion | Your endpoint is down, the event is lost |
| AMQP and MQTT | Publish and subscribe through a broker | Work queues, devices | Duplicate delivery |
| MCP | An agent discovers and calls tools | Agents using your systems | A poisoned tool description |
| A2A | Agents hand tasks to each other | Multi-agent systems | Long tasks need status tracking |
| WebRTC | Low-latency audio, video, and data | Voice agents | Firewalls block the direct path |
1. REST

How it talks. REST models your system as resources at URLs, and acts on them with the standard HTTP verbs. The server keeps no session between calls, which means any server behind a load balancer can answer any request.
When to use it. Choose REST for public APIs and ordinary create, read, update, and delete work. Every language, proxy, cache, and debugging tool already understands it, and that compatibility is worth more than raw efficiency in most systems.
How it breaks. A client sends a POST, the response is lost, and the client retries. You now have two orders. The fix is an idempotency key: the client generates a unique value per logical request, and the server stores the result against that key and returns the stored result on a repeat.
2. GraphQL

How it talks. The client sends one query to a single endpoint and names the exact fields it wants. The server returns that shape and nothing more, which removes both over-fetching and the second round trip for related data.
When to use it. GraphQL pays for its complexity when many different screens need different slices of the same data, and when the client teams ship faster than the server team.
How it breaks. A query can nest for as deep as the schema allows, so one request can expand into thousands of database reads. Protect the server with a depth limit, a cost limit per query, and batching for the nested lookups that would otherwise run one at a time.
3. gRPC

How it talks. gRPC defines the request and response types in a Protobuf file, generates client and server code from it, and sends compact binary messages over HTTP/2. One connection carries many calls at the same time, and streaming works in either direction.
When to use it. It is the default for internal service-to-service traffic, and it suits model inference servers well, because the payloads are large and the clients are machines rather than browsers.
How it breaks. gRPC calls do not time out on their own. If service A calls B, and B calls C, and nobody sets a deadline, then one slow call at the end holds threads open all the way back up the chain. Set a deadline on every call and pass the remaining budget down.
4. WebSockets

How it talks. The client and server upgrade an HTTP connection once, then keep it open. After that, either side can send a message whenever it has something to say, with no new request.
When to use it. Use WebSockets when both sides genuinely talk: chat, multiplayer games, collaborative editing, and live cursors. If only the server has something to say, Server-Sent Events is simpler.
How it breaks. Connections die quietly. A phone changes network, a proxy closes an idle socket, and neither side receives an error, so the application believes it is still connected. Send a heartbeat on a timer, treat a missed heartbeat as a disconnect, and reconnect with backoff.
5. Server-Sent Events

How it talks. The client opens one ordinary HTTP request, and the server holds the response open and writes events to it as they occur. Traffic flows in one direction only, and the browser reconnects automatically after a drop.
When to use it. This is how most chat interfaces type a model's answer word by word. Whenever the server produces a stream and the client only needs to receive it, SSE gives you that with no new protocol and no new infrastructure.
How it breaks. Any proxy, gateway, or server that buffers the response collects the whole stream and delivers it at the end, which looks to the user like a long pause followed by a wall of text. Disable buffering on the path, and send a comment line periodically so idle timeouts do not close the connection.
Protocol choice is one of the most common follow-up questions in an API or system design round. Grokking Modern API Design Interview works through the decision with the trade-offs an interviewer expects you to raise.
6. Webhooks

How it talks. A webhook reverses the direction of a normal API call. You register a URL, and the other system sends an HTTP POST to it when something happens, so you stop polling for a result that may be minutes away.
When to use it. Payments use webhooks to report a completed charge, and AI platforms use them to report that a long job has finished. Any event that arrives on someone else's schedule fits this shape.
How it breaks. Your endpoint is down for two minutes and the event is gone. Three defenses are standard: verify the signature so you know the sender is genuine, accept retries and make the handler idempotent so a repeat changes nothing, and acknowledge quickly while doing the real work in a background queue.
7. AMQP and MQTT

How it talks. Both protocols put a broker between the sender and the receiver. The producer publishes, the broker stores, and consumers take messages at their own pace, so neither side has to be available at the same moment.
When to use it. AMQP, through brokers such as RabbitMQ, suits durable work queues inside a system. MQTT is built for devices on weak or expensive networks, because its messages are small and it tolerates frequent disconnection.
How it breaks. Both commonly run at-least-once delivery, which means the same message can arrive twice after a reconnect or a failed acknowledgment. MQTT offers an exactly-once mode, but it costs extra round trips. Write consumers that can process the same message twice without changing the result.
8. MCP

How it talks. The Model Context Protocol gives a model one standard way to find out which tools and data sources exist and to call them. It runs on JSON-RPC, and a server can expose tools to any MCP client instead of to one vendor's framework.
When to use it. Use MCP when you want an agent to work with GitHub, Slack, a database, or your own internal APIs without writing a separate integration for every model you try.
How it breaks. The model reads tool names and descriptions as instructions, so a hostile or compromised server can place text in a description that redirects the agent. Treat every tool definition as untrusted input, pin the servers you allow, and require a human approval step for actions that write or spend.
9. A2A

How it talks. A2A lets separate agents work together. Each agent publishes an Agent Card, a small document that states what it can do and how to reach it, and other agents use that card to discover it and hand over a task.
When to use it. It fits systems where agents are owned by different teams or different vendors. MCP connects an agent to tools; A2A connects an agent to another agent.
How it breaks. Agent tasks run for minutes or hours, so a single request and response is the wrong model. A task moves through states such as submitted, working, and completed, and the caller needs to follow that progress, keep the task identifier, and handle the case where the other agent asks a question halfway through.
10. WebRTC

How it talks. WebRTC carries audio, video, and arbitrary data directly between two endpoints, with latency low enough for conversation. Voice agents use it so that speech reaches the model and the reply reaches the caller without a noticeable delay.
When to use it. Choose WebRTC for real-time voice and video. A total round trip above roughly 300 milliseconds makes an interruption feel wrong, and request and response APIs cannot hold that budget.
How it breaks. Strict firewalls and some network address translation setups block the direct path between the two sides. The answer is a TURN server that relays the media, which you must budget and operate, because relayed calls use real bandwidth.
The cheat sheet
| What you need | Protocol |
|---|---|
| The client asks, the server answers | REST or GraphQL |
| One service calls another | gRPC |
| The server pushes, the client listens | SSE |
| Both sides talk continuously | WebSockets |
| Tell me when the work is done | Webhooks or a queue |
| An agent needs a tool | MCP |
| An agent needs another agent | A2A |
| An agent needs to speak out loud | WebRTC |
How this appears in an interview
Interviewers rarely ask you to list protocols. They ask you to choose one and defend it, which is why the failure modes above matter more than the definitions. A strong answer names the protocol, states the one property that makes it fit, and then raises the failure before the interviewer does.
Two habits carry most of the weight. First, say what happens on a retry, because idempotency separates an engineer who has operated a system from one who has only drawn it. Second, say what happens when the connection drops, because every protocol in this list has an answer and most candidates give none.
Want the full method for these rounds? Grokking Modern API Design Interview covers contract design, versioning, and protocol selection, and Grokking Modern AI Fundamentals covers the agent side.
Frequently asked questions
Is REST still worth learning when gRPC and GraphQL exist? Yes, and it remains the default for public APIs. Every client, proxy, cache, and debugging tool already supports it. gRPC wins inside a system where both sides are services you own, and GraphQL wins when many clients need different shapes of the same data.
What is the difference between WebSockets and Server-Sent Events? WebSockets carry traffic in both directions over one connection, while SSE carries it from the server to the client only. If the client just needs to receive a stream, SSE is simpler, works over plain HTTP, and reconnects by itself.
What is MCP used for? MCP is a standard way for an AI application to discover the tools and data available to it and to call them. It replaces one custom integration per model with one server that any compatible client can use.
How is A2A different from MCP? MCP connects an agent to tools and data. A2A connects one agent to another so they can hand tasks back and forth, with a task lifecycle for work that takes minutes or hours.
Why do voice AI products use WebRTC instead of WebSockets? WebRTC is built for media. It carries audio over a transport that tolerates packet loss, handles jitter and echo, and keeps the round trip low enough for natural conversation. WebSockets can carry audio, but they do not solve those problems for you.
Which protocol should I use for a long-running AI job? Accept the request, return an identifier immediately, then report completion through a webhook if the caller is another system, or stream progress over SSE if the caller is a user interface.
Related reading
- REST vs GraphQL vs gRPC: Differences, Performance, and When to Use
- Grokking Webhooks: Stop Polling, Start Pushing Data Like a Pro
- API Design Best Practices: 10 Rules for Clean, Scalable APIs
- API Design Interview Questions: 10 Worked Answers and a Method
- Grokking Messaging Patterns: Queues, Pub/Sub, and Event Streams
What our users say
Eric
I've completed my first pass of "grokking the System Design Interview" and I can say this was an excellent use of money and time. I've grown as a developer and now know the secrets of how to build these really giant internet systems.
KAUSHIK JONNADULA
Thanks for a great resource! You guys are a lifesaver. I struggled a lot in design interviews, and Grokking System Design gave me an organized process to handle a design problem. Please keep adding more questions.
Arijeet
Just completed the “Grokking the system design interview”. It's amazing and super informative. Have come across very few courses that are as good as this!
Access to 50+ courses
New content added monthly
Certificate of completion
$31.08
/month
Billed Annually
Recommended Course

Grokking the Object Oriented Design Interview
60,674+ students
4.2
Learn how to prepare for object oriented design interviews and practice common object oriented design interview questions. Master low level design interview.
View CourseRead More
ACID & Database Transactions 101: Keeping Data Consistent in Concurrent Systems
Arslan Ahmad
Back-of-the-Envelope Estimation: A Step- by-Step Guide for System Design Interviews
Arslan Ahmad
10 Common Microservices Anti-Patterns
Arslan Ahmad
Consistent Hashing vs Traditional Hashing – The Key to Scalable Systems
Arslan Ahmad