Skip to content

How Kubernetes Works Internally: API Server, etcd, and the Watch Mechanism

What really happens when you run kubectl apply? Follow one Deployment through the API server, etcd, watches, informers, controllers, scheduler and kubelet.

Prefer to watch? This post is the written companion to the video above.

You type one line:

kubectl apply -f deployment.yaml

A few seconds later, three containers are running on machines you never touched. No component received a direct order to start them. Nobody called the scheduler. Nobody called the node.

So what actually happened?

The short answer is that Kubernetes is not a system of components calling each other. It is a shared database of desired state, plus a set of independent programs that each watch that database and nudge the real world until it matches. Once that idea clicks, everything else in Kubernetes stops looking like magic.

Let's follow that one command all the way down.

The big picture

Here are the main players. Keep this map in mind; we will walk through each box.

fig. 1 / cluster architecture
Kubernetes architecture kubectl sends requests to kube-apiserver. kube-apiserver is the only component that reads and writes etcd. The controller manager, the scheduler and the kubelet on each worker node all watch the API server and write changes back through it. The kubelet drives the container runtime through CRI. control plane worker node https only path to etcd watch + write kubectl you kube-apiserver front door + watches etcd source of truth controller-manager reconcile loops scheduler picks a node kubelet node agent containerd via CRI request / watch event read / write to etcd
Every component talks to the API server. Only the API server talks to etcd.

Two rules explain most of this diagram:

  1. Everyone talks to the API server. kubectl, the controllers, the scheduler and every kubelet are all clients of kube-apiserver.
  2. Only the API server talks to etcd. No other component reads or writes the database directly.

The API server: the only door into the cluster

kube-apiserver is a REST API over HTTPS. Every Kubernetes object, whether a Pod, a Deployment or a Secret, is a resource at a URL like:

/apis/apps/v1/namespaces/default/deployments/web

When kubectl sends your Deployment, the request passes through a fixed pipeline before anything is stored:

  1. Authentication. Who are you? A client certificate, a bearer token, an OIDC token and so on.
  2. Authorization. Are you allowed to do this? Usually RBAC: can this user create deployments in default?
  3. Mutating admission. Plugins and webhooks may change the object, for example to fill in defaults or inject a sidecar.
  4. Schema validation. Is this a well-formed Deployment?
  5. Validating admission. Plugins and webhooks may reject the object, but can no longer change it.
  6. Persist to etcd.

By default, kubectl apply does a client-side apply: it compares your file with the live object and the last-applied-configuration annotation, then sends either a create or a patch. With --server-side, the API server does that merge itself and tracks which tool owns which field. Either way, the request lands in the same pipeline.

Why put a single API server in front of everything? Because it gives the cluster one place to enforce security, validation and versioning, and one place to serve change notifications. That second job is the one that makes Kubernetes work, and we will get to it in a moment.

etcd: the source of truth

etcd is a small, strongly consistent, distributed key-value store. It uses the Raft consensus algorithm, so a write is only acknowledged once a majority of etcd members have it. That is why production clusters run 3 or 5 etcd members.

Kubernetes stores every object under a key that mirrors its path:

/registry/deployments/default/web
/registry/replicasets/default/web-7d9c6b8f5
/registry/pods/default/web-7d9c6b8f5-x2k4q

Built-in objects are stored in a compact binary encoding (protobuf), not as YAML. Custom resources are stored as JSON. You never touch those bytes yourself; the API server converts between versions and encodings for you.

The most important detail about etcd for this story is the revision. etcd keeps a single counter for the whole store, and every write bumps it. Each key remembers the revision at which it was last modified. Kubernetes exposes that number on every object as metadata.resourceVersion.

metadata:
  name: web
  namespace: default
  resourceVersion: "184213"

Treat resourceVersion as an opaque string, not a number you do math on. Kubernetes uses it in two ways:

  • Optimistic concurrency. An update must carry the resourceVersion it was based on. If someone else changed the object in the meantime, the write fails with a 409 Conflict and the client retries with fresh data. No locks needed.
  • Watching. "Tell me everything that changed after version 184213." That is the heart of the next section.

The watch mechanism

Polling would not scale. Imagine thousands of kubelets and dozens of controllers asking "anything new?" every second. Instead, clients open a watch: a long-lived HTTP request that streams events as they happen.

kubectl get pods --watch -v=6
# ... GET /api/v1/namespaces/default/pods?resourceVersion=184213&watch=true

Each event is small and typed:

{ "type": "ADDED",    "object": { "kind": "Pod", "metadata": { "name": "web-x2k4q", "resourceVersion": "184220" } } }
{ "type": "MODIFIED", "object": { "kind": "Pod", "metadata": { "name": "web-x2k4q", "resourceVersion": "184231" } } }

The types are ADDED, MODIFIED, DELETED, plus BOOKMARK (a heartbeat that just says "you are up to date as of this resourceVersion") and ERROR.

The standard pattern every client follows is list, then watch:

  1. LIST the resources to get the current state and the collection's resourceVersion.
  2. WATCH starting from that resourceVersion, so no change between the two calls is missed.
  3. If the connection drops, resume the watch from the last resourceVersion you saw.
  4. If the server answers 410 Gone (your version is too old, history was compacted), start over with a fresh LIST.

Recent releases refine step 1. With the watch list feature (beta, on by default since Kubernetes 1.34 on the server and 1.35 in client-go), an informer can skip the big LIST: it opens a watch with sendInitialEvents=true, receives the current state as a stream of ADDED events, then a BOOKMARK that says "you are caught up". Same idea, much less memory on the API server for large clusters.

One more trick keeps this cheap. The API server does not open an etcd watch per client. It keeps an in-memory watch cache for each resource type, fed by its own watch on etcd, and fans events out to all of its clients from there. A thousand kubelets watching pods do not mean a thousand watches on etcd.

Informers: the local cache every component keeps

Writing correct list-then-watch code, with reconnects and 410 handling, is fiddly. So Kubernetes components (and almost every operator you will ever install) use a library pattern from client-go called an informer.

fig. 2 / inside an informer
Watch and informer flow The API server streams watch events to a Reflector, which lists once and then watches. Changes go into a DeltaFIFO queue. Each change updates the Indexer, the local cache, and calls event handlers. Handlers enqueue the object's key into a workqueue. A reconcile worker takes keys from the queue, reads the current object from the local cache and writes any changes back to the API server. events enqueue key read (Lister) write: create / update / delete API server watch cache Reflector LIST, then WATCH DeltaFIFO ordered changes Indexer local cache of objects event handlers add / update / delete workqueue "namespace/name" keys reconcile() worker
A controller never polls. It lists once, watches forever, keeps a local cache, and works from a queue of keys.

The pieces, in order:

  • Reflector. Does the LIST and WATCH against the API server and handles reconnects.
  • DeltaFIFO. A queue of changes ("deltas") in the order they happened.
  • Indexer. A thread-safe, in-memory copy of every object of that type. This is the component's local cache.
  • Event handlers. Callbacks for add, update and delete. In a controller, they do almost nothing: they just put the object's key, like default/web, onto a workqueue.
  • Workqueue. Deduplicates keys and rate-limits retries. If a Deployment changes five times in a second, the key is processed once with the latest state.

This design has a big consequence: controllers read from their local cache, not from the API server. When a controller asks "which pods belong to this ReplicaSet?", it answers from memory using a Lister. Reads are free; only writes go over the network.

It also means the cache can be slightly stale. Controllers are written to tolerate that, which is the next idea.

Controllers and the reconcile loop

A controller is a loop that compares desired state (the spec you wrote) with actual state (what exists, usually reflected in status and in other objects) and takes one step to close the gap.

fig. 3 / the reconcile loop

yes

no

observe
read desired + actual
from the cache

do they match?

do nothing

act
create, update or delete
via the API server

the change produces
new watch events

Observe, compare, act. The result of acting is just another change to observe.

In Go-flavored pseudocode, the core of the ReplicaSet controller looks like this:

func (c *ReplicaSetController) reconcile(key string) error {
    rs, err := c.rsLister.Get(key)            // from the local cache
    if notFound(err) {
        return nil                            // it was deleted, nothing to do
    }
    pods := c.podLister.Owned(rs)             // pods with an ownerReference to rs
    diff := len(pods) - int(*rs.Spec.Replicas)

    switch {
    case diff < 0:
        return c.createPods(rs, -diff)        // too few: create more
    case diff > 0:
        return c.deletePods(rs, pickVictims(pods, diff)) // too many: delete some
    }
    return c.updateStatus(rs, pods)
}

This is simplified on purpose. The real ReplicaSet controller also tracks expectations ("I just asked for 2 pods, wait until I see them in my cache") so a slightly stale cache never makes it create duplicates, and it creates pods in growing batches (1, 2, 4, 8 and so on) so a broken pod template fails fast instead of flooding the API server.

Notice what the loop does not do. It never asks "what event just happened?" It looks at the whole current state and decides what to do. This is called being level-triggered rather than edge-triggered. If the controller crashes, misses an event, or restarts, the next reconcile still does the right thing, because it only cares about where things are now.

kube-controller-manager is one binary that runs dozens of these loops: Deployments, ReplicaSets, Jobs, Nodes, EndpointSlices, ServiceAccounts and more. Each one owns a small slice of the cluster's behavior, and none of them call each other.

The scheduler: binding pods to nodes

The scheduler is just another controller with a very specific job. It watches for pods whose spec.nodeName is empty and answers one question: which node should run this pod?

For each pending pod it runs two phases:

  1. Filter. Drop nodes that cannot run the pod: not enough free CPU or memory requests, a taint the pod does not tolerate, a node selector or affinity that does not match, a port already in use, and so on.
  2. Score. Rank the remaining nodes, for example to spread replicas across zones or to prefer nodes that already have the image.

Then it binds the pod: it posts a Binding to the API server, which sets spec.nodeName on the pod. That is all. The scheduler does not start anything. It writes one field and moves on.

The kubelet: making it real on the node

Each node runs a kubelet. It watches the API server for pods whose spec.nodeName equals its own name, so it only ever hears about its own work.

When a new pod shows up for its node, the kubelet talks to the container runtime (containerd or CRI-O) over the Container Runtime Interface (CRI), a gRPC API:

  1. RunPodSandbox: create the pod's sandbox, including its network namespace. The runtime calls the CNI plugin to give the pod an IP address.
  2. Pull the image if it is not already on the node.
  3. CreateContainer and StartContainer for each container (init containers first).

Then the kubelet does the other half of its job: it reports back. It updates the pod's status through the API server (phase, IP, container states, readiness), runs liveness and readiness probes, and keeps doing this for as long as the pod lives.

Again, the kubelet does not take orders. It reconciles: desired pods for this node versus containers actually running.

One Deployment, from kubectl to running container

Now let's put it all together. Here is the full journey of a Deployment named web with replicas: 3.

fig. 4 / one Deployment, end to end
kubeletschedulerrs ctrldeploy ctrletcdapiserverkubectlapply Deployment1authn, authz, admission2write Deployment3Deployment ADDED4create ReplicaSet5write ReplicaSet6ReplicaSet ADDED7create 3 Pods8write Pods9Pods ADDED, no node10bind pod to node-111write nodeName12Pod bound to node-113CRI: sandbox, pull, start14status: Running, Ready15write status16
Solid arrows are requests and writes. Dashed arrows are watch events streamed by the API server.

Step by step:

  1. kubectl sends the Deployment to the API server.
  2. The API server authenticates, authorizes, runs admission and validation, then writes the object to etcd. The Deployment now exists, but nothing is running.
  3. The Deployment controller gets an ADDED event through its informer. It reconciles: "this Deployment wants a ReplicaSet for this pod template, and none exists." It creates a ReplicaSet named after the Deployment plus a hash of the pod template.
  4. The ReplicaSet controller sees the new ReplicaSet. Desired 3, actual 0. It creates 3 Pods, each with an ownerReference pointing back at the ReplicaSet.
  5. The scheduler sees 3 pods with no nodeName. For each one it filters and scores nodes, then writes a Binding.
  6. The kubelet on each chosen node sees a pod assigned to it. It asks the runtime to create the sandbox, set up networking, pull the image and start the containers.
  7. The kubelet reports status. When the pods become Ready, the ReplicaSet and Deployment controllers update their own status fields, and if a Service selects these pods, the EndpointSlice controller adds their IPs so traffic can reach them.

Every arrow in that sequence is either a write to the API server or a watch event from it. No component ever calls another one directly. That is the whole trick.

And it is also why Kubernetes heals itself. Delete one of those pods by hand, and the ReplicaSet controller sees "desired 3, actual 2" and creates a new one. The same loop that built the system keeps it alive.

Summary

  • The API server is the only door. All reads and writes go through it, and it is the only component that talks to etcd.
  • etcd is the source of truth. Every write gets a new revision, exposed as resourceVersion.
  • Watches replace polling. Clients list once, then stream changes from a resourceVersion, served from the API server's watch cache.
  • Informers give every component a local cache and a queue of keys to work on, so reads are cheap.
  • Controllers reconcile desired against actual state, level-triggered, so missed events and restarts are harmless.
  • The scheduler only writes nodeName. The kubelet makes it real through CRI and reports status back.

Next time you run kubectl apply, picture the packets: one write into etcd, then a ripple of watch events, each component doing one small job, until the real world matches what you asked for.