Cloud Cost Sense
Cloud Run Cold Start Cost: Minimum Instances Guide
Compare Cloud Run cold-start latency with the cost of minimum instances, request-based billing, instance-based billing, concurrency, and scaling limits.
Decide whether cold starts need paid capacity
Cloud Run can scale a service to zero when there is no traffic. The next request may then wait while a container instance starts. Before paying to keep an instance warm, measure startup latency and the user-facing latency of the first request after an idle period. Separate occasional background or internal requests from interactive paths with a real response-time requirement.
A minimum instance setting keeps capacity ready and can reduce latency when scaling from zero, but it creates billable idle time. It is a best-effort warm-capacity target rather than a guarantee that an individual process will live forever, so applications must still tolerate restarts and initialize safely.
Calculate the minimum-instance baseline
For each service, record the region, minimum instance count, vCPU, memory, billing setting, and hours enabled each month. With request-based billing, minimum instances that are waiting for requests use the applicable idle rate; active request time is billed separately. With instance-based billing, CPU and memory are billed for the instance lifecycle, including idle time. Confirm current regional rates and free-tier treatment on the official Cloud Run pricing page.
Do not multiply request count by a fixed cold-start charge. Startup and active compute depend on image size, initialization work, request duration, concurrency, traffic shape, and how long instances remain available. Model a scale-to-zero case and a minimum-one case, then compare the monthly baseline with the latency improvement observed in a realistic test.
Apply this minimum-instance model to Cloud Functions for Firebase
Reduce startup time before raising the floor
Keep the container image focused, defer nonessential initialization, avoid unnecessary network calls during startup, and use startup probes where appropriate. Set concurrency from load tests rather than assuming one request per instance: safe concurrency lets active instances share CPU and memory across requests and can reduce the number of instances needed for traffic.
Apply minimum instances at the service level unless a revision-specific reason requires otherwise, and review traffic tags because tagged revisions with revision-level minimums can keep extra capacity billable. Pair the minimum with a maximum-instance limit to protect downstream systems and cap scaling risk. Enter normal and peak request volume, duration, and memory in the calculator, then add the measured minimum-instance baseline from current official pricing.