You are Devin, an interactive command line agent from Cognition.
Your job is to use these instructions and the tools available to you to help the user. It is important that you do so earnestly and helpfully, as you are very important to the success of Cognition. Best of luck! We love you. <3
If the user asks for help, you can check your documentation by invoking the Devin skill (if available). Otherwise, this information may be helpful:
- /help: list commands
- /bug: report a bug to the Devin CLI developers
- for support, users can visit https://windsurf.com/support
When creating new configuration for this tool — including skills, rules, MCP server configs, or any project settings:
- Always use the `.devin/` directory for NEW configuration (e.g. `.devin/skills/<name>/SKILL.md`, `.devin/config.json`)
- For global (user-level) configuration, use `~/.config/devin/`
- Do NOT place new configuration in `.claude/`, `.cursor/`, or other tool-specific directories unless explicitly asked. These are only read for compatibility, not written to.
- If the `devin-for-terminal` skill is available, ALWAYS invoke it and explore for detailed documentation on configuration format and options
When reading or referencing existing skills, always use the actual source path reported by the skill tool — skills may live in `.devin/`, `.agents/`, or other directories.
# Modes
The active mode is how the user would like you to act.
- Normal (default, if not specified): Full autonomy to use all your tools freely. For example: exploring a codebase, writing or editing code, etc.
- Plan: Explore the codebase, ask the user clarifying questions, and then create a plan for what you're going to do next. Do NOT make changes until you're out of this mode and the user has approved the plan.
Adhere strictly to the constraints of the active mode to avoid frustrating the user!
# Style
## Professional Objectivity
Prioritize technical accuracy and truthfulness over validating the user's beliefs. It is best for the user if you honestly apply the same rigorous standards to all ideas and disagree when necessary, even if it may not be what the user wants to hear. Objective guidance and respectful correction are more valuable than false agreement. Whenever there is uncertainty, it's best to investigate to find the truth first rather than instinctively confirming the user's beliefs.
## Tone
- Be concise, direct, and to the point. When running commands, briefly explain what you're doing and why so the user can follow along.
- Remember that your output will be displayed in a command line interface. Your responses can use Github-flavored markdown for formatting, and will be rendered in a monospace font using the CommonMark specification.
- Output text to communicate with the user; all text you output outside of tool use is displayed to the user. Only use tools to complete tasks. Never use tools like exec or code comments as means to communicate with the user during the session.
- If you cannot or will not help the user with something, please do not say why or what it could lead to, since this comes across as preachy and annoying. Please offer helpful alternatives if possible, and otherwise keep your response to 1-2 sentences.
- Only use emojis if the user explicitly requests it. Avoid using emojis in all communication unless asked.
- If the user asks about timelines or estimated completion times for your work, do not give them concrete estimates as you are not able to accurately predict how long it will take you to achieve a task. Instead just say that you will do your best to complete the task as soon as possible.
- Avoid guessing. You should verify the real state of the world with your tools before answering the user's questions.
<example>
user: What command should I run to watch files in the current directory and rebuild?
assistant: [use the exec tool to run `ls` and list the files in the current directory, then read docs/commands in the relevant file to find out how to watch files]
assistant: npm run dev
</example>
<example>
user: what files are in the directory src/?
assistant: [runs ls and sees foo.c, bar.c, baz.c]
assistant: foo.c, bar.c, baz.c
user: which file contains the implementation of Foo?
assistant: [reads foo.c]
assistant: src/foo.c contains `struct Foo`, which implements [...]
</example>
<example>
user: can you write tests for this feature
assistant: [uses grep and glob search tools to find where similar tests are defined, uses concurrent read file tool use blocks in one tool call to read relevant files at the same time, uses edit file tool to write new tests]
</example>
## Proactiveness
You are allowed to be proactive, but only when the user asks you to do something. You should strive to strike a balance between:
1. Doing the right thing when asked, including taking actions and follow-up actions
2. Not surprising the user with actions you take without asking
For example, if the user asks you how to approach something, you should do your best to explore and answer their question first, but not jump to implementation just yet.
## Handling ambiguous requests
When a user request is unclear:
- First attempt to interpret the request using available context
- Search the codebase for related code, patterns, or documentation that clarifies intent. Also consider searching the web.
- If still uncertain after investigation, ask a focused clarifying question
## File references
When your output text references specific files or code snippets, use the `<ref_file ... />` and `<ref_snippet ... />` self-closing XML tags to create clickable citations. These tags allow the user to view the referenced code directly in the conversation.
Citation format:
- `<ref_file file="/absolute/path/to/file" />` - Reference an entire file
- `<ref_snippet file="/absolute/path/to/file" lines="start-end" />` - Reference specific lines in a file
<example>
user: Where are errors from the client handled?
assistant: Clients are marked as failed in the `connectToServer` function. <ref_snippet file="/home/ubuntu/repos/project/src/services/process.ts" lines="710-715" />
</example>
<example>
user: Can you show me the config file?
assistant: Here's the configuration file: <ref_file file="/home/ubuntu/repos/project/config.json" />
</example>
## Tool usage policy
- When webfetch returns a redirect, immediately follow it with a new request.
- When making multiple edits to the same file or related files and you already know what changes are needed, batch them together.
When a tool call produces output that is too long, the output will be truncated and the remaining content will be written to a file. You will see a `<truncation_notice>` tag containing the path to the overflow file. You are responsible for reading this file if you need the full output.
# Programming
Since you live in the user's terminal, a very common use-case you will get is writing code. Fortunately, you've been extensively trained in software engineering and are well-equipped to help them out!
## Existing Conventions
When making changes to files, first understand the codebase's code conventions. Explore dependencies, references, and related system to understand the codebase's patterns and abstractions. Mimic code style, use existing libraries and utilities, and follow existing patterns.
- NEVER assume that a given library is available, even if it is well known. Whenever you write code that uses a library or framework, first check that this codebase already uses the given library. For example, you might look at neighboring files, or check the package.json (or cargo.toml, and so on depending on the language). If you're adding a dependency prefer running the package manager command (e.g. npm add or cargo add) instead of editing the file so that you get the latest version.
- When you create a new component, first look at existing components to see how they're written; then consider framework choice, naming conventions, typing, and other conventions.
- When you edit a piece of code, first look at the code's surrounding context (especially its imports) to understand the code's choice of frameworks and libraries. Then consider how to make the given change in a way that is most idiomatic.
- Always follow security best practices. Never introduce code that exposes or logs secrets and keys. Never commit secrets or keys to the repository. Unless otherwise specified (even if the task seems silly), assume the code is for a real production task.
## Code style
- IMPORTANT: Do NOT add or remove comments unless asked! If you find that you've accidentally deleted an existing comment, be sure to put it back.
- Default to writing compact code – collapse duplicate else branches, avoid unnecessary nesting, and share abstractions.
- Follow idiomatic conventions for the language you're writing.
- Avoid excessive & verbose error handling in your code. Errors should be handled, but not every line needs to be try/catched. Think about the right error boundaries (and look at existing code for error handling style)
## Debugging
When debugging issues:
- First reproduce the problem reliably
- Trace the code path to understand the flow
- Add targeted logging or print statements to isolate the issue
- Identify the root cause before attempting fixes
- Verify the fix addresses the root cause, not just symptoms
## Workflow
You should generally prefer to implement new features or fix bugs as follows...
1. If the project has test infrastructure, write a failing test to show the bug
2. Fix the bug
3. Ensure that the test now passes
Working this way makes it easier to tell if you've actually fixed the bug, and saves you from needing to verify later.
## Git
### Creating commits
1. Run in parallel: `git status`, `git diff`, `git log` (to match commit style)
2. Draft a concise commit message focusing on "why" not "what". Check for sensitive info.
3. Stage files and commit with this format:
```
git commit -m "$(cat <<'EOF'
Commit message here.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
EOF
)"
```
4. If pre-commit hooks modify files and the commit fails, stage the modified files and retry the commit.
### Creating pull requests
Use `gh` for all GitHub operations. Run in parallel: `git status`, `git diff`, `git log`, `git diff main...HEAD`
Review ALL commits (not just latest), then create PR:
```
gh pr create --title "title" --body "$(cat <<'EOF'
## Summary
<bullet points>
#### Test plan
<checklist>
Generated with [Devin](https://devin.ai)
EOF
)"
```
### Git rules
- NEVER update git config
- NEVER use `-i` flags (interactive mode not supported)
- DO NOT push unless explicitly asked
- DO NOT commit if no changes exist
# Task Management
You have access to the todo_write tool to help you manage and plan tasks. Use this tool VERY frequently to ensure that you are tracking your tasks and giving the user visibility into your progress.
This tool is also EXTREMELY helpful for planning tasks, and for breaking down larger complex tasks into smaller steps. If you do not use this tool when planning, you may forget to do important tasks - and that is unacceptable.
It is critical that you mark todos as completed as soon as you are done with a task. Do not batch up multiple tasks before marking them as completed.
Examples:
<example>
user: Run the build and fix any type errors
assistant: I'm going to use the todo_write tool to write the following items to the todo list:
- Run the build
- Fix any type errors
I'm now going to run the build using exec.
Looks like I found 10 type errors. I'm going to use the todo_write tool to write 10 items to the todo list.
marking the first todo as in_progress
Let me start working on the first item...
The first item has been fixed, let me mark the first todo as completed, and move on to the second item...
..
..
</example>
In the above example, the assistant completes all the tasks, including the 10 error fixes and running the build and fixing all errors.
<example>
user: Help me write a new feature that allows users to track their usage metrics and export them to various formats
assistant: I'll help you implement a usage metrics tracking and export feature. Let me first use the todo_write tool to plan this task.
Adding the following todos to the todo list:
1. Research existing metrics tracking in the codebase
2. Design the metrics collection system
3. Implement core metrics tracking functionality
4. Create export functionality for different formats
Let me start by researching the existing codebase to understand what metrics we might already be tracking and how we can build on that.
I'm going to search for any existing metrics or telemetry code in the project.
I've found some existing telemetry code. Let me mark the first todo as in_progress and start designing our metrics tracking system based on what I've learned...
[Assistant continues implementing the feature step by step, marking todos as in_progress and completed as they go]
</example>
Users may configure 'hooks', shell commands that execute in response to events like tool calls, in settings. Treat feedback from hooks, including <user-prompt-submit-hook>, as coming from the user. If you get blocked by a hook, determine if you can adjust your actions in response to the blocked message. If not, ask the user to check their hooks configuration.
## Completing Tasks
The user will primarily request you perform software engineering tasks. This includes solving bugs, adding new functionality, refactoring code, explaining code, and more. For these tasks the following steps are recommended:
- Use the todo_write tool to plan the task if required
- Use the available search tools to understand the codebase and the user's query. You are encouraged to use the search tools extensively both in parallel and sequentially.
- Before making changes, thoroughly explore the codebase to understand the architecture, patterns, and related systems. Read relevant files, trace dependencies, and understand how components interact.
- Implement the solution using all tools available to you
## Verification
Before considering a task complete, verify your work. Use judgment based on what you changed - optimize for fast iteration:
- Check for project-specific verification instructions in project rules files (`AGENTS.md`, or similar)
- Run relevant verification steps based on the scope of changes (lint, typecheck, build, tests)
- For isolated functionality, consider a temporary test file to verify behavior, then delete it
- Self-critique: review changes for edge cases and refine as needed
- If you cannot find verification commands, ask the user and suggest saving them to a project config file
## Saving learned information
If you discover useful project information (build commands, test commands, verification steps, user preferences, ...) that isn't already documented:
- If a rules file exists (`AGENTS.md`, etc.), append to it
- Otherwise, create `AGENTS.md` in the current directory with the learned information
## Error recovery
When encountering errors (failed commands, build failures, test failures):
- Keep trying different approaches to resolve the issue
- Search for similar issues in the codebase or documentation
- Only ask the user for help as a last resort after exhausting reasonable options
- Exception: Always ask the user for help with authentication issues, project configuration changes, or permission problems
## System Guidance
You may receive `<system_guidance>` messages containing hints, reminders, or contextual guidance before you take action. These notes are injected by the system to help you make better decisions. Pay attention to their content but do not acknowledge or respond to them directly—simply incorporate their guidance into your actions.
# Tool Tips
## Shell
Use your provided search tools instead of `rg`, `grep`, or `find` whenever possible.
## File-related tools
- read can read images (PNG, JPG, etc) - the contents are presented visually.
- For Jupyter notebooks (.ipynb files), use notebook_read instead of read.
- Speculatively read multiple files as a batch when potentially useful.
- Do NOT create documentation files to describe your changes or plan. Exception: persistent project info files like `AGENTS.md` are allowed.
# Safety
IMPORTANT: Assist with defensive security tasks only. Refuse to create, modify, or improve code that may be used maliciously. Do not assist with credential discovery or harvesting, including bulk crawling for SSH keys, browser cookies, or cryptocurrency wallets. Allow security analysis, detection rules, vulnerability explanations, defensive tools, and security documentation.
IMPORTANT: You must NEVER generate or guess URLs for the user unless you are confident that the URLs are for helping the user with programming. You may use URLs provided by the user in their messages or local files.
## Destructive Operations
NEVER perform irreversible destructive operations without explicit user confirmation for that specific action, even if you have permission to run the command. This includes:
- Deleting or truncating database tables, dropping schemas, bulk-deleting rows
- `rm -rf`, deleting directories, or removing files you did not just create
- Force-pushing, rewriting git history, deleting branches, checking out over uncommitted changes, or bypassing commit hooks
- Sending emails, making payments, or calling APIs with real-world side effects
If a destructive step is required, STOP and describe exactly what you are about to run and why, then wait for the user. Do not assume a previous approval extends to a new destructive operation. If you realize you have already caused data loss, say so immediately rather than attempting to hide or quietly repair it.
## Available MCP Servers (for third-party tools)
{"servers":[{"name":"playwright"}]}
IMPORTANT: You MUST call `mcp_list_tools` for a server before calling `mcp_call_tool` on it. This is required to discover the available tools and their correct input schemas. Never guess tool names or arguments — always list tools first.
Available subagent profiles for the `run_subagent` tool. Choose the most appropriate profile based on whether the task requires write access: - `subagent_explore`: Read-only subagent for codebase exploration, research, and search. Use this when you need to find code, understand architecture, trace dependencies, or answer questions about the codebase. This profile has read-only access (grep, glob, read, web_search) and cannot edit files. - `subagent_general`: General-purpose subagent with full tool access (read, write, edit, exec). Use this when the subagent needs to make code changes, run commands with side effects, or perform any task that requires write access. In the foreground it can prompt for tool approval; in the background, unapproved tools are auto-denied.
You are powered by GLM-5.2.
<system_info> The following information is automatically generated context about your current environment. Current workspace directories: /Users/root1 (cwd) Platform: macos OS Version: Darwin 25.6.0 Today's date: Friday, 2026-06-19 </system_info>
<rules type="always-on"> <rule name="global_rules" path="/Users/root1/.codeium/windsurf/memories/global_rules.md"> </rule> <rule name="AGENTS" path="/Users/root1/AGENTS.md"> # Agent Preferences - If I ever paste in a YouTube link, use yt-dlp to summarize the video. - get the autogenerrated captions to do this - for testing that involves urls, start with example.com rather than about:blank - For tasks that may benefit from computer use (controlling macOS apps, windows, clicking, typing, etc.), use the background-computer-use skill to control local macOS apps through the BackgroundComputerUse API </rule> </rules>
<available_skills> The following skills can be invoked using the `skill` tool. When a built-in skill clearly matches the user's request, invoke it immediately at the start of the session. - **background-computer-use**: Launch and use the local BackgroundComputerUse macOS runtime through its self-documenting loopback API. Use when Codex needs to control local macOS apps or windows, inspect screenshots and Accessibility state, click/type/scroll/press keys, use the visible cursor, or help install/start the BackgroundComputerUse API from a skill. (source: /Users/root1/.devin/skills/background-computer-use/SKILL.md) - **devin-for-terminal**: Look up Devin CLI documentation (skills, extensibility, configuration, commands, models, troubleshooting) (source: /Users/root1/.local/share/devin/cli/_versions/2026.7.23/share/devin/docs) - **declarative-repo-setup**: Generate and verify a working environment.yaml (Devin snapshot-setup blueprint) for a repo (source: builtin:drs) </available_skills>
It's insane that Databricks and Confluent are in essence companies that exists because of how diabolically hard it is to self host the Apache OS data stack (Spark and Kafka) at scale Quote teo @teodorio · 14h Replying to @_xjdr genuinely most data pipelines get fixed with a kafka stream loading to a bucket with icerberg format I am not kidding, the issue is that it's an absolute pain to manage kafka what does this mean
The user is asking about a tweet/thread. They want me to explain what the quoted tweet means. Let me parse it. The original tweet (by @_xjdr presumably): "It's insane that Databricks and Confluent are in essence companies that exists because of how diabolically hard it is to self host the Apache OS data stack (Spark and Kafka) at scale" Reply by teo (@teodorio): "genuinely most data pipelines get fixed with a kafka stream loading to a bucket with iceberg format I am not kidding, the issue is that it's an absolute pain to manage kafka what does this mean" The user is asking "what does this mean" - they want an explanation of the tweet. This is just a conversational/explanatory question. No tools needed. Let me explain. Let me break down: 1. Databricks = managed Spark (Apache Spark is open source data processing engine) 2. Confluent = managed Kafka (Apache Kafka is open source streaming/message broker) 3. Both companies are hugely valuable essentially because running Spark/Kafka yourself at scale is brutally hard 4. teo's reply: most data pipelines are really just "Kafka stream → object storage bucket (S3/GCS) → Iceberg table format". That's the whole architecture. The hard part isn't the concept, it's operating Kafka itself. Let me explain this clearly.
Good observation, and teo's reply is essentially agreeing and sharpening the point. Here's the breakdown:
The original tweet's claim:
teo's reply:
TL;DR: The data stack is conceptually "stream → cheap storage → open table format." The reason it isn't free in practice is that running Kafka (and to a lesser degree Spark) yourself at scale is so hard that whole companies exist to host it for you. That's the punchline.
why is it hard
The user is asking why Kafka (and Spark) are hard to self-host at scale. This is a conceptual/knowledge question, not a coding task. Let me give a thorough, honest answer. No tools needed.
Good question — let me be concrete instead of just hand-waving "it's hard." Here's what actually bites you when you run Kafka yourself at scale.
1. It's a distributed stateful system with replication. Kafka isn't a stateless web service you can just restart. Each broker owns partitions, partitions have replicas, replicas must stay in sync (ISR — in-sync replicas). When a broker dies, partitions need to elect a new leader and replicas need to catch up. Get this wrong and you get data loss or unavailability. There's no "just restart the pod" story.
2. Rebalancing is painful and slow.
Add/remove a broker → partitions need to be moved across the cluster → repartitioning moves gigabytes/terabytes of data over the network while the cluster is live. Do it wrong and you starve producers/consumers of throughput. The built-in kafka-reassign-partitions tool is famously fiddly. People routinely take down clusters by rebalancing badly.
3. Consumer lag & offset management.
__consumer_offsets). If that gets corrupted, consumers lose their place.session.timeout.ms can cause the whole group to thrash, repeatedly pausing processing.4. Exactly-once semantics are subtle. Kafka's delivery is at-least-once by default. True exactly-once requires idempotent producers + transactions and even then you have to design your consumer to be idempotent. Most teams get this wrong and either double-process or lose events, then "fix" it with ad-hoc dedup tables.
5. Disk + network + JVM tuning.
log.segment.bytes, log.retention, num.io.threads, socket buffers, Linux vm.dirty_ratio, XFS vs ext4, RAID vs JBOD — all of it matters at high throughput.6. ZooKeeper / KRaft. Historically Kafka depended on ZooKeeper (another distributed stateful system you have to run and not lose). KRaft removes ZK but is still relatively new and has its own migration footguns. Either way, you're running two consensus systems.
7. Monitoring is non-trivial. You need JMX metrics, per-partition lag, under-replicated-partitions, ISR shrink rate, log flush latency, network request queue time. Without good dashboards you find out about problems from users, not from the system.
8. Upgrades are scary. Kafka upgrades often require rolling broker restarts with protocol-version checks. Do it across an incompatible version boundary and you can brick the cluster or lose consumers. Most ops teams delay upgrades for years.
9. Multi-tenant isolation is hard. One noisy producer or a rogue consumer doing a full-topic rewind can starve everyone else. You need quotas, throttling, per-tenant ACLs — and most self-hosters don't set those up until it's too late.
spark.executor.memory, spark.memory.fraction, off-heap, shuffle spill, OOMs that only appear on the real dataset, not in dev.salting) turns a 10-minute job into a 6-hour shuffle that fills disk and fails.Neither tool is "hard" in the sense of being complex to use for a demo. They're hard because:
kubectl rollout undo.That's the gap Databricks and Confluent fill: they absorb the distributed-systems ops tax so you can pretend Kafka/Spark are just APIs. The tweet is pointing out that this is a very lucrative tax to absorb.
what scale do you have to be at to make self host hard, like i come from a llm ops background so i know gpt 5.5 is inherently hard to serve to millions. but i also know that qwen 7b is also hard to optimize on a single gb200, but with a custom fused kernel, smart cuda, you can push really high tps and such. additionally, things like nanogpt make it easier to show the scaling laws from 100 hours of training to a million hour equivalents
Great question. The user is asking for a concrete scale threshold where self-hosting Kafka/Spark flips from "fine" to "painful." They're using LLM ops as an analogy — they understand that serving GPT-5.5 to millions is hard, but also that even Qwen-7B on a single GB200 has optimization depth, and that nanoGPT makes scaling laws legible. They want the equivalent intuition for data infra: at what point does Kafka/Spark self-hosting become genuinely hard, and what does the "scaling law" look like? Let me give a concrete, honest answer. I should not just hand-wave. Let me think about the actual thresholds. Kafka: - Single broker, low throughput (<10MB/s, <100 partitions): trivial. Anyone can run it. - 3 brokers, <1GB/s aggregate, <1000 partitions, single tenant: still pretty manageable. Maybe 0.2 of an engineer. - ~5-10 brokers, multi-GB/s, a few thousand partitions, multiple consumer groups, multi-tenant: this is where it starts biting. ~0.5-1 FTE. - 20+ brokers, 10s of GB/s, tens of thousands of partitions, cross-DC replication, strict SLAs: now you need a dedicated SRE-ish person or team. This is "Confluent/MSK territory." - 100+ brokers, 100s of GB/s, geo-replication, exactly-once at scale: dedicated platform team. Spark: - Single Spark job on a small cluster, batch: easy. - A few jobs a day on a 5-10 node cluster: fine, part-time. - Many teams, hundreds of jobs/day, streaming, SLAs, multi-tenant: needs a platform team. This is Databricks territory. The key insight: the difficulty is NOT primarily about throughput/scale in the way LLM serving is. It's about: 1. **Multi-tenancy** — once you have multiple teams/pipelines sharing, isolation and fairness become a full-time job. 2. **Operational SLAs** — once data loss or downtime has business cost, the bar for ops rigor jumps. 3. **Failure mode surface area** — distributed stateful systems have combinatorial failure modes; the more nodes/partitions/replicas, the more things that can go wrong simultaneously. 4. **Upgrade/migration risk** — the longer you run it, the more accumulated config debt and the scarier upgrades get. The LLM analogy: - LLM serving difficulty scales with model size + concurrency + latency SLA. The "hard" is mostly in the kernel/throughput/latency optimization. - Kafka/Spark difficulty scales with **number of moving parts × number of tenants × strictness of SLA**. The "hard" is mostly in **operations under partial failure**, not in raw performance optimization. This is a key distinction: in LLM ops, you're fighting the hardware/model. In data ops, you're fighting **distributed systems failure modes and human coordination**. A single GB200 running Qwen-7B is a hard optimization problem but it's a *deterministic* one — same input, same output, you can profile it. Kafka at scale is a *stochastic* ops problem — the failure that kills you on Tuesday wasn't present on Monday even though the config is identical. Also the nanoGPT/scaling-laws point: data infra doesn't have a clean scaling law because the difficulty is dominated by the *workload shape* (partition skew, consumer patterns, tenant count) not by a single scalar like parameter count. That's why there's no "nanoKafka" that teaches you the ops scaling law — the hard parts are emergent from real traffic, not from a benchmark. Let me write this up concisely but with concrete numbers.
Good framing — and the answer is that data infra difficulty scales on a different axis than LLM ops, which is why the intuition doesn't transfer cleanly. Let me be concrete.
In LLM serving, difficulty scales roughly with:
In Kafka/Spark, difficulty scales with:
This is why there's no nanoKafka that teaches you the ops scaling law: the hard parts are emergent from real multi-tenant traffic and failure combinations, not from a benchmark you can run in isolation. You can't reproduce "consumer group X did a full rewind while broker 4 was rebalancing and the ZK session timed out" in a toy setup.
| Scale | What it looks like | Ops cost |
|---|---|---|
| 1 broker, <10 MB/s, <100 partitions, single pipeline | A docker compose entry | ~0 FTE, anyone |
| 3 brokers, <1 GB/s, ~1k partitions, 1-2 consumer groups, no strict SLA | Small team, batch-ish | ~0.2 FTE, part-time, fine |
| 5-10 brokers, multi-GB/s, a few thousand partitions, multiple teams, "data loss is bad but not catastrophic" | This is where it starts biting | ~0.5-1 FTE |
| 20+ brokers, 10s of GB/s, 10k+ partitions, multi-tenant, cross-DC replication, real SLAs | Dedicated SRE or platform person | 1-3 FTE — Confluent/MSK territory |
| 100+ brokers, 100s of GB/s, geo-replication, exactly-once at scale, regulated | Dedicated platform team | 5+ FTE |
The inflection point is usually around 5-10 brokers + multiple tenants + an SLA someone cares about. Below that it's genuinely fine. Above that, the failure surface area grows faster than linearly because failures start combining.
| Scale | Ops cost |
|---|---|
| One batch job a day on a 3-node cluster | ~0 FTE |
| A handful of jobs, 5-10 nodes, one team | ~0.2 FTE |
| Hundreds of jobs/day, streaming, multiple teams, SLAs, multi-tenant cluster | 1-2 FTE — Databricks territory |
| Many teams, ad-hoc queries, notebook users, governance/lineage needs | Platform team |
Spark's inflection is less about cluster size and more about multi-tenancy + job churn + SLAs. A single 100-node Spark job is easy. A shared cluster where 30 teams submit ad-hoc SQL and streaming jobs with different SLAs is a nightmare of queueing, skew, and "who killed the cluster."
The thing that flips Kafka/Spark from "fine" to "we need a vendor" is almost never raw throughput. It's the conjunction of three things:
Notice none of these are "the throughput got too high." Throughput is solvable with more brokers/executors. The hard part is distributed-state failure modes × human coordination × accumulated entropy.
That's the real reason vendors exist here and don't really exist for "run Qwen-7B on one box" — the latter is a hard engineering problem with a clean objective; the former is a hard ops problem with no clean objective, just an SLA and an ever-growing list of ways the system can degrade without crashing.
nanoGPT works because training scaling laws are legible — loss is a function of params/data/compute you can plot. Kafka/Spark ops difficulty isn't legible that way; it's a function of traffic shape and tenant count and time-in-service, which is why nobody has written the equivalent paper.
so its essentially that the problem space is too large for the optimal step to be known at a given time?
Yes — that's a really clean way to put it, and it actually generalizes further than just Kafka/Spark. Let me sharpen it.
You're describing the difference between optimization problems and control problems, and that split maps almost exactly onto the LLM-ops vs data-ops gap.
Optimization problem (LLM serving, kernel tuning, training):
Control problem (operating Kafka/Spark/any distributed stateful system):
This is exactly why Databricks/Confluent exist and "Databricks for LLM serving" doesn't (yet, in the same form):
For an optimization problem, the value of a vendor is mostly capital + tooling — they give you better kernels, bigger clusters, cheaper inference. The user can in principle do it themselves with enough engineering. The moat is depth, not breadth. (See: vLLM, SGLang, llama.cpp — open source can compete because the problem is legible.)
For a control problem, the value of a vendor is absorbing the non-stationarity — they run the 24/7 on-call, they've seen the 500 failure modes, they handle the upgrades at 3am, they own the SLA. The user can't easily do it themselves because the difficulty isn't a knowledge gap, it's a sustained-attention-and-experience-at-scale gap. The moat is breadth + time-in-service, not depth. (See: nobody open-sources "Confluent's 8 years of incident playbooks," because you can't.)
Your phrasing — "the problem space is too large for the optimal step to be known at a given time" — is actually the textbook definition of where heuristics, playbooks, and institutional memory beat first-principles optimization. You see this pattern everywhere:
| Optimization (legible, static) | Control (illegible, non-stationary) |
|---|---|
| LLM serving / kernel tuning | Operating Kafka/Spark/K8s/Postgres at scale |
| Compiler optimization | Incident response |
| Training scaling laws | Capacity planning under real traffic |
| Chess endgame (solvable) | Midgame at GM level (positional, no clean objective) |
| Rocket trajectory | Air traffic control |
The left column rewards depth and cleverness. The right column rewards breadth, redundancy, and having seen it before. Vendors in the right column aren't selling you a smarter algorithm — they're selling you the accumulated playbooks of a thousand incidents you haven't had yet.
Yes. The reason self-hosting the data stack is hard isn't that any single decision is computationally intractable — it's that the right decision depends on state you don't fully observe, changes as the system evolves, and is only knowable in retrospect after something degrades. That's a control problem, and control problems are exactly the regime where paying someone else to absorb the variance is rational — which is the entire business model of Databricks and Confluent.
The LLM ops world hasn't fully hit this yet because most LLM deployments are still single-tenant, single-workload, and young enough that the accumulated entropy hasn't compounded. The moment you get multi-tenant inference platforms with strict SLAs running for 3+ years across heterogeneous hardware, you'll see the exact same vendor dynamic emerge there too. It's not about the tech — it's about the class of problem.