What to Expect in the Modal System Design Interview

Expect a design question based on Modal's own infrastructure. Modal is a serverless platform that runs Python functions in containers on CPUs and GPUs. One candidate reports a design task in the process, though its format is not public, and Modal does not publish a question list.

Its engineering posts do describe the problems it solves every day, which are container starts in under a second, fast loading of large model files, and scheduling scarce GPUs. Prepare for questions close to those problems, because a generic social network question is unlikely here.

The Question Types

A serverless function platform. Design a system where a user submits a Python function and the platform runs it on demand. The hard parts are the cold start, autoscaling, and isolation between customers. A cold start is the delay when a new container must boot before it handles its first request. Modal's docs split that delay into two parts: time waiting in a queue for a container, and setup work inside the new container.

Fast container image loading. Design storage so that a container can start without downloading its whole image. Modal's public approach loads images lazily, so the image becomes a small index of file metadata, about five megabytes, that mounts in a few milliseconds. File contents are fetched on demand from a chain of caches, which runs from local memory to local disk, then a zone cache server, a regional content delivery network, and blob storage (large object storage in the cloud).

GPU scheduling. Design a scheduler that places jobs on machines with the right GPU type, in the right region, across several clouds. GPUs are scarce and unevenly placed, so the scheduler must balance wait time against cost, and you should expect questions about queues, priorities, and what happens when no GPU is free.

Usage metering and billing. Modal charges only for the seconds that code runs, so design the pipeline that records usage per second, per customer, without losing or double counting events.

What the Interviewer Grades

Practical judgment matters more than unusual parts, so state requirements first, including the latency target and the failure cases. Name your trade-offs out loud, such as cost against cold start time, and connect each choice to the customer, because a slow start is a slow product for every user of the platform. Show that you know where time is spent in a container start, because candidates who can name the stages of a boot score well.

A Walkthrough: Design a Serverless GPU Function Platform

Here is a high level plan for the signature question.

1. Requirements (5 minutes). Many customers, each with private code and data. A user defines a function and its dependencies, and the platform must run it on request, scale to zero when idle, and start new containers fast, with a target such as most cold starts under one second.

2. Image build and storage. Build each customer's container image once, and store it as an index plus content-addressed file chunks. Content-addressed means each chunk is named by a hash of its bytes, so identical files are stored once.

3. Lazy loading at start. When a container starts, mount the index only and fetch file contents on first read through the cache chain. Cache the most common files, such as Python packages, on every worker in advance.

4. The scheduler. Keep a pool of workers across regions and clouds, and match each request to a worker with free capacity of the right GPU type. Queue requests when no worker fits, start new workers when the queue grows, and stop idle containers after a set idle period.

5. Warm capacity and snapshots. Keep a small number of warm containers (already booted and waiting) for latency sensitive functions, and use memory snapshots for the rest. A memory snapshot saves the state of a booted container, so later starts restore it instead of repeating setup.

6. Isolation and metering. Run each customer's code in an isolated container so one customer cannot read another's data. Record start and stop events for every container, then sum them per customer for billing.

Common Mistakes in This Round

  • Starting with Kubernetes. Naming a tool is not a design, so explain what the scheduler must do and then say which parts you would build or buy.
  • Ignoring the cold start. A design that downloads a full image on every start fails the core requirement.
  • No answer for scarce GPUs. Say what happens when demand exceeds the GPU pool: queueing, priorities, or a different region.
  • Forgetting isolation. Customer code is untrusted, so mention the isolation layer early, before the interviewer asks.

How to Prepare

TAGS
System Design Interview
CONTRIBUTOR
Arslan Ahmad
Arslan Ahmad
ex-FAANG engineering manager and author or Grokking series.

GET YOUR FREE

Coding Questions Catalog

Design Gurus Newsletter - Latest from our Blog
Boost your coding skills with our essential coding questions catalog.
Take a step towards a better tech career now!
Explore Answers
What to Expect in the Plaid System Design Interview
The design topics Plaid asks about, how they map to its bank data network, and one example question solved step by step.
What are Datadog interview questions?
What is the main goal of Microsoft?
What is AEO and How to Make Answers LLM-Friendly?
Learn what Answer Engine Optimization (AEO) is, how to write LLM-friendly answers, and how it helps your content rank higher on Google, ChatGPT, Perplexity, and Gemini.
What is exit code 14 in MongoDB?
How to Introduce Yourself in a Mock Interview (+ Example)
Introduce Yourself in a Mock Interview: 30-Second Template
Related Courses
New
Grokking the AI System Design Interview course cover
Grokking the AI System Design Interview
Learn to design AI systems the way interviewers expect: classic ML products, LLM and RAG architectures, and agentic systems, all through the lens of the system design interview.
4.6
(3,192 learners)
Discounted price for Your Region

$99

Grokking the Coding Interview: Patterns for Coding Questions course cover
Grokking the Coding Interview: Patterns for Coding Questions
The 24 essential patterns behind every coding interview question. Available in Java, Python, JavaScript, C++, C#, and Go. The most comprehensive coding interview course with 543 lessons. A smarter alternative to grinding LeetCode.
4.6
Discounted price for Your Region

$197

Grokking Modern AI Fundamentals course cover
Grokking Modern AI Fundamentals
Master the fundamentals of AI today to lead the tech revolution of tomorrow.
4.1
Discounted price for Your Region

$72

Design Gurus logo
One-Stop Portal For Tech Interviews.
Copyright © 2026 Design Gurus, LLC. All rights reserved.