[{"data":1,"prerenderedAt":1063},["ShallowReactive",2],{"blog-alternates-\u002Fen\u002Fschema-change-30-second-alarm-billion-row-night":3,"blog-nav-facets-en":6,"blog-post-en-schema-change-30-second-alarm-billion-row-night":11},{"en":4,"pt":5},"\u002Fen\u002Fschema-change-30-second-alarm-billion-row-night","\u002Fmudanca-schema-alarm-30s-noite-bilhao-linhas",{"populatedCategories":7,"authorCount":10},[8,9],"engenharia","dooh",1,{"post":12,"authorIndex":1024,"counterpartSlug":1030,"related":1031},{"id":13,"title":14,"authors":15,"body":17,"category":8,"cover":990,"description":993,"discussionUrl":994,"draft":995,"extension":996,"featured":995,"highlights":997,"meta":1008,"navigation":1009,"ogImage":991,"path":1010,"publishedAt":1011,"seo":1012,"slug":1013,"sourceHash":994,"stem":1014,"tags":1015,"translate":1009,"translatedFrom":1020,"translationKey":1021,"translationStatus":1022,"updatedAt":994,"__hash__":1023},"posts_en\u002Fen\u002Fposts\u002Fmudanca-schema-alarm-30s-noite-bilhao-linhas.md","A schema change, a 30-second alarm, and a billion-row night",[16],"paulo-ramires",{"type":18,"value":19,"toc":957},"minimark",[20,24,29,33,36,39,58,68,78,81,84,87,91,94,97,100,106,109,125,128,131,134,137,140,143,146,150,153,156,159,167,170,173,190,193,196,199,203,210,218,222,225,228,232,237,240,254,257,260,263,267,270,273,276,282,285,288,291,295,298,301,304,307,310,313,316,322,325,330,333,337,340,343,346,349,352,355,358,361,364,367,370,374,377,382,385,390,393,396,401,404,407,412,415,418,423,426,430,433,436,439,443,447,450,453,456,459,462,466,469,489,492,497,501,504,507,510,514,517,520,534,539,543,549,552,556,559,562,565,568,571,574,580,583,586,589,592,597,605,608,612,615,635,638,641,645,650,653,656,661,664,667,670,675,678,681,686,689,692,695,698,703,706,709,712,716,719,722,739,742,745,748,758,761,765,768,771,776,779,782,787,790,793,796,799,802,805,808,811,815,818,821,835,838,841,845,848,870,875,882,885,889,892,898,908,914,917,921,924,927,930,933,936,939,942,945,948,951,954],[21,22],"phase-marker",{"tone":23},"incident",[25,26,28],"h2",{"id":27},"the-short-version","The short version",[30,31,32],"p",{},"On 2–3 September 2026, Lucimark production entered a Durable Object SQLite storage storm.",[30,34,35],{},"We had just shipped a Proof of Play schema change: storing player location as a geohash trail on 15-minute rollups instead of baking geography into each row's unique key. The data model was sound. The migration strategy was not sized for our largest tenant — hundreds of thousands of rollup rows inside a single TenantDO.",[30,37,38],{},"Two failure modes, in sequence.",[30,40,41,45,46,53,54,57],{},[42,43,44],"strong",{},"Day one — writes."," A synchronous rebuild during Durable Object startup produced roughly 625 million billed SQLite writes within a few hours. At ",[47,48,52],"a",{"href":49,"rel":50},"https:\u002F\u002Fdevelopers.cloudflare.com\u002Fdurable-objects\u002Fplatform\u002Fpricing\u002F",[51],"nofollow","Cloudflare's Durable Objects SQLite list prices"," (docs last updated 25 August 2026), Workers Paid writes cost $1.00 \u002F million rows and reads $0.001 \u002F million — about 1,000×. That is a gross list-price equivalent ",[42,55,56],{},"before"," included allowances, not the invoiced amount.",[30,59,60,63,64,67],{},[42,61,62],{},"Night and next morning — reads."," We moved the remaining work into an alarm-driven asynchronous backfill — a batched replay of what boot could not finish. That stopped the startup failures, but with the UNIQUE index down, the state machine repeatedly executed O(n) SQL — time proportional to the number of rows in the table — at roughly 30-second cadence. Hourly read peaks reached approximately 0.3–1.4 billion billed rows. Reads were much cheaper per row. The danger was different: ",[42,65,66],{},"the migration had become an unbounded loop."," Nothing in the platform would naturally stop it.",[69,70],"blog-figure",{":breakout":71,":height":72,":n":73,":width":74,"alt":75,"caption":76,"src":77},"true","941","1","1672","Comparison of Cloudflare Durable Objects SQLite list prices for writes and reads against the volumes observed during the incident.","Writes caused the immediate cost shock. Reads were about 1,000× cheaper per row at list price, but exposed an unbounded loop. Gross list-price equivalents before included allowances; not the invoice.","\u002Fimg\u002Fposts\u002Fdo-storage-storm\u002Fdurable-objects-reads-vs-writes-cost.png",[30,79,80],{},"This was not a GraphQL artifact, a billing glitch, or runaway Proof of Play ingest. Ingest that day was only thousands of rollup inserts. Worker HTTP traffic had increased modestly.",[30,82,83],{},"Once we understood the loop, containment was direct: park alarm-driven housekeeping, deploy cooldown and a helper index, stop re-arming heavy states every 30 seconds, remove disposable historical Proof of Play rollups from the affected tenant, and mark the migration complete. Reads fell from tens of millions per hour toward baseline the same morning.",[30,85,86],{},"The storm was in the data plane — cost and saturation, not a fleet-wide playback outage.",[25,88,90],{"id":89},"why-this-architecture-made-the-incident-possible","Why this architecture made the incident possible",[30,92,93],{},"Lucimark is a DOOH platform for operating screen networks from the browser.",[30,95,96],{},"Our API runs on Cloudflare Workers. Each customer tenant is represented by a Durable Object — a TenantDO — backed by its own SQLite database.",[30,98,99],{},"Proof of Play rollups live there too: 15-minute and daily summaries describing what played, where, and when.",[69,101],{":breakout":71,":height":72,":n":102,":width":74,"alt":103,"caption":104,"src":105},"2","Lucimark incident architecture showing API, isolated tenant Durable Object, SQLite Proof of Play storage and unaffected screen playback.","The incident stayed inside that tenant's Proof of Play SQLite. The screen fleet was not the blast radius.","\u002Fimg\u002Fposts\u002Fdo-storage-storm\u002Flucimark-tenantdo-incident-architecture.png",[30,107,108],{},"That architecture gives us useful properties:",[110,111,112,116,119,122],"ul",{},[113,114,115],"li",{},"tenant isolation;",[113,117,118],{},"data locality;",[113,120,121],{},"a clear blast radius;",[113,123,124],{},"no shared relational database becoming the bottleneck for every customer.",[30,126,127],{},"But it also creates a billing surface.",[30,129,130],{},"Durable Object SQLite charges for rows read and rows written. Writes are dramatically more expensive per row than reads, but either can become operationally significant when an algorithm touches an entire large table repeatedly.",[30,132,133],{},"A high-volume tenant in this architecture is not simply \"more HTTP.\"",[30,135,136],{},"It is more rows inside one SQLite database, potentially being revisited by one state machine every time that Durable Object wakes.",[30,138,139],{},"Median tenants crossed this migration without incident.",[30,141,142],{},"One large rollup table did not.",[30,144,145],{},"Cloudflare provides budget notifications, but they do not stop runtime consumption. The system itself still needs engineering bounds.",[25,147,149],{"id":148},"what-we-were-trying-to-ship","What we were trying to ship",[30,151,152],{},"Our Proof of Play 15-minute rows originally carried geography inside their natural key — the business key that identifies the rollup, distinct from the internal primary key.",[30,154,155],{},"That meant geography was effectively part of record identity: changing location information could fork what should otherwise represent the same logical rollup.",[30,157,158],{},"The new model stored a geohash trail separately and removed geography from the unique key.",[69,160],{":breakout":71,":height":161,":n":162,":width":163,"alt":164,"caption":165,"src":166,":hideOnMobile":71},"887","3","1774","Before and after Proof of Play schema showing geo removed from the natural_key and stored as geohash_trail.","Geography left rollup identity and moved into the geohash_trail column. Technical identifiers match the schema.","\u002Fimg\u002Fposts\u002Fdo-storage-storm\u002Fproof-of-play-geohash-schema-before-after.png",[168,169],"schema-compare",{},[30,171,172],{},"That required the migration to:",[110,174,175,178,181,184,187],{},[113,176,177],{},"introduce the new data representation;",[113,179,180],{},"update existing rollup rows;",[113,182,183],{},"change the natural key;",[113,185,186],{},"deduplicate collisions;",[113,188,189],{},"rebuild uniqueness.",[30,191,192],{},"For a small tenant, this is a migration.",[30,194,195],{},"For our largest affected tenant — roughly 540k 15-minute rollups plus 280k daily rows — it was effectively a rewrite of one of the hottest historical tables in the object.",[30,197,198],{},"We underestimated that table.",[25,200,202],{"id":201},"timeline","Timeline",[30,204,205,206,209],{},"Times are ",[42,207,208],{},"BRT (UTC−3)",", the timezone used in the operational post-mortem.",[69,211],{":breakout":71,":height":212,":n":213,":width":214,"alt":215,"caption":216,"src":217,":hideOnMobile":71},"724","4","2172","Timeline of the Lucimark Durable Objects incident from deployment on September 2 through recovery on September 3.","From deploy to recovery, in two windows: a write burst on 2 September, then an alarm-driven read storm. Overview only; the table below is the readable source.","\u002Fimg\u002Fposts\u002Fdo-storage-storm\u002Fdurable-objects-incident-timeline.png",[219,220],"incident-timeline",{":items":221},"[{\"when\":\"2 Sep ~14:16 BRT\",\"event\":\"geohash-trail migration deploy\",\"result\":\"Heavy work on Durable Object startup.\"},{\"when\":\"2 Sep 14:00–15:00 BRT\",\"event\":\"Write burst\",\"result\":\"Two consecutive hourly buckets, ~333M and ~290M billed rows written\u002Fh — we do not assign which hour to which total. 503s and resets on the large tenant.\"},{\"when\":\"2 Sep late afternoon\",\"event\":\"Migration leaves boot\",\"result\":\"Alarm-driven async backfill. The 503s stop; the amplifier only moves.\"},{\"when\":\"2 Sep overnight\",\"event\":\"Read storm\",\"result\":\"Writes normalize. Billed reads ~0.3–1.4B\u002Fh while UNIQUE is down.\"},{\"when\":\"3 Sep morning\",\"event\":\"Alarms parked (~09:20) and hotfix\",\"result\":\"The scan loop stops. Residual still ~57M–74M reads\u002Fh.\"},{\"when\":\"3 Sep ~12:01 BRT\",\"event\":\"Historical PoP removed\",\"result\":\"Non-authoritative operational dataset; forward ingest continues. Migration marked done.\"}]",[21,223],{"tone":224},"recovery",[30,226,227],{},"The hourly read curve after that belongs in the recovery section — not here.",[25,229,231],{"id":230},"root-cause-four-things-that-became-dangerous-together","Root cause: four things that became dangerous together",[233,234,236],"h3",{"id":235},"_1-heavy-work-on-a-large-table","1. Heavy work on a large table",[30,238,239],{},"The migration required operations equivalent to:",[110,241,242,245,248,251],{},[113,243,244],{},"temporarily remove uniqueness;",[113,246,247],{},"update existing rows;",[113,249,250],{},"deduplicate;",[113,252,253],{},"recreate the UNIQUE index.",[30,255,256],{},"While the UNIQUE structure was unavailable, operations such as deduplication and index reconstruction had to touch large portions of the table.",[30,258,259],{},"On hundreds of thousands of rows, that matters.",[30,261,262],{},"Repeatedly, it becomes the incident.",[233,264,266],{"id":265},"_2-moving-sync-work-to-async-did-not-make-it-safe","2. Moving sync work to async did not make it safe",[30,268,269],{},"Our first implementation did too much work during Durable Object startup.",[30,271,272],{},"That caused resets and 503s.",[30,274,275],{},"Moving the work to an alarm-driven state machine was the correct response to the startup problem.",[277,278,279],"pull-quote",{},[30,280,281],{},"Async ≠ safe.",[30,283,284],{},"The synchronous version had a natural failure boundary: startup could time out or the object could reset.",[30,286,287],{},"The asynchronous version could keep waking indefinitely.",[30,289,290],{},"Moving expensive work off the request path removed one operational constraint without adding another.",[233,292,294],{"id":293},"_3-alarm-cadence-became-a-cost-multiplier","3. Alarm cadence became a cost multiplier",[30,296,297],{},"During the migration, states roughly looked like:",[30,299,300],{},"pending → updating → deduping → indexing",[30,302,303],{},"The batched row updates themselves were manageable.",[30,305,306],{},"The dangerous states were the ones that could require O(n) work — scanning or rebuilding the whole table.",[30,308,309],{},"The backfill could re-arm itself at roughly 30-second cadence while still incomplete.",[30,311,312],{},"On a table containing more than half a million rollup rows, that meant a theoretically expensive operation could be revisited around 120 times per hour.",[30,314,315],{},"This is where a slow migration becomes a runaway one.",[69,317],{":breakout":71,":height":72,":n":318,":width":74,"alt":319,"caption":320,"src":321,":hideOnMobile":71},"5","Thirty-second migration amplifier loop showing update, dedupe, unique index creation, failure and retry around 120 times per hour.","A ~30 s alarm turned O(n) migration work into a repeating amplifier. The ×120\u002Fh follows from the cadence, not from a separate measurement.","\u002Fimg\u002Fposts\u002Fdo-storage-storm\u002Fdurable-objects-30-second-amplifier-loop.png",[323,324],"amplifier-loop",{},[30,326,327],{},[42,328,329],{},"Alarm cadence is a cost multiplier.",[30,331,332],{},"Scheduling frequency cannot be treated as an implementation detail when the scheduled operation scales with table size.",[233,334,336],{"id":335},"_4-failed-unique-creation-had-no-breaker","4. Failed UNIQUE creation had no breaker",[30,338,339],{},"Duplicates remained.",[30,341,342],{},"The UNIQUE index rebuild failed.",[30,344,345],{},"The error was logged.",[30,347,348],{},"Then the state machine tried again.",[30,350,351],{},"And again.",[30,353,354],{},"We found roughly 2,800 migration tick failures, all associated with the same large tenant.",[30,356,357],{},"The worst possible shape is straightforward:",[30,359,360],{},"scan → fail → wait 30 seconds → scan again",[30,362,363],{},"A failed DDL operation in a recurring state machine is not merely an error message.",[30,365,366],{},"It is a loop.",[30,368,369],{},"And loops need breakers.",[25,371,373],{"id":372},"what-we-ruled-out","What we ruled out",[30,375,376],{},"Several plausible explanations did not survive the data.",[30,378,379],{},[42,380,381],{},"GraphQL or billing artifact",[30,383,384],{},"Namespace-scoped Durable Object metrics matched the log behavior and the sharp recovery after removing the historical dataset.",[30,386,387],{},[42,388,389],{},"Proof of Play ingest explosion",[30,391,392],{},"Ingest volume was only thousands of rollup inserts that day.",[30,394,395],{},"That could not explain hundreds of millions or billions of rows being touched.",[30,397,398],{},[42,399,400],{},"Account-wide Worker traffic",[30,402,403],{},"HTTP traffic increased roughly 2–3×.",[30,405,406],{},"The SQLite movement was orders of magnitude larger.",[30,408,409],{},[42,410,411],{},"Every tenant failing",[30,413,414],{},"Other production tenants did not show the same backfill-failure signature.",[30,416,417],{},"The problem was concentrated in one large table.",[30,419,420],{},[42,421,422],{},"Frontend outage",[30,424,425],{},"The 503s on day one came from TenantDO migration resets — wakes that needed the object during the synchronous rebuild. After the migration moved async, the dominant issue became housekeeping churn, not a blank screen across the fleet.",[25,427,429],{"id":428},"product-impact","Product impact",[30,431,432],{},"On the affected tenant, during the synchronous migration, some operations that needed to wake its TenantDO — including parts of the Screens experience — could fail with 503s. Other tenants did not show the same signature.",[30,434,435],{},"For that high-volume tenant, we deliberately removed historical Proof of Play rollups involved in the migration. This was a dataset-specific operational decision: those rollups were not authoritative contractual evidence and forward ingest continued normally.",[30,437,438],{},"Deleting data is not a generic migration strategy. In this incident, continuing an unbounded rewrite of replaceable history was worse than rebuilding history forward from a clean state.",[25,440,442],{"id":441},"how-we-mitigated-it","How we mitigated it",[233,444,446],{"id":445},"first-stop-the-multiplier","First: stop the multiplier",[30,448,449],{},"The migration was driven by TenantDO alarms.",[30,451,452],{},"So the first useful lever was not shutting down Proof of Play ingest.",[30,454,455],{},"It was parking alarm-driven housekeeping.",[30,457,458],{},"That stopped the repeating scan loop.",[30,460,461],{},"This distinction is now part of our operational model: the correct kill switch must correspond to the amplifier causing the incident.",[233,463,465],{"id":464},"then-fix-the-state-machine","Then: fix the state machine",[30,467,468],{},"The production hotfix introduced:",[110,470,471,474,477,480,483,486],{},[113,472,473],{},"approximately 15-minute cooldowns around heavy dedupe\u002Findex states;",[113,475,476],{},"a non-unique helper index while uniqueness is temporarily unavailable;",[113,478,479],{},"short-circuiting when the desired UNIQUE index is already present;",[113,481,482],{},"no immediate UNIQUE retry while duplicates still exist;",[113,484,485],{},"failed index creation returning to dedupe instead of retrying the same full-table operation;",[113,487,488],{},"migration progress decoupled from the normal 30-second active alarm cadence.",[30,490,491],{},"The key rule is simple:",[277,493,494],{},[30,495,496],{},"Never treat O(n) SQL like a heartbeat.",[233,498,500],{"id":499},"then-eliminate-unnecessary-migration-work","Then: eliminate unnecessary migration work",[30,502,503],{},"Even with cooldown and helper indexing, rebuilding the historical dataset in place would still have been expensive.",[30,505,506],{},"The durable fix for this tenant was to eliminate the need to migrate those disposable historical rows at all.",[30,508,509],{},"Once the historical rollups were removed and the migration state marked complete, the scan surface disappeared.",[233,511,513],{"id":512},"finally-verify-recovery","Finally: verify recovery",[30,515,516],{},"We did not call the incident over because a deployment succeeded.",[30,518,519],{},"We looked for the actual signals:",[110,521,522,525,528,531],{},[113,523,524],{},"rows read per hour fell sharply;",[113,526,527],{},"migration tick failures went to zero;",[113,529,530],{},"TenantDO behavior returned to baseline;",[113,532,533],{},"normal housekeeping could be restored.",[277,535,536],{},[30,537,538],{},"The recovery curve mattered more than the deploy status.",[540,541],"recovery-sequence",{":points":542},"[\"74M\",\"2.9M\",\"836k\",\"424k\"]",[69,544],{":breakout":71,":height":72,":n":545,":width":74,"alt":546,"caption":547,"src":548,":hideOnMobile":71},"6","Logarithmic recovery chart showing Durable Objects SQLite reads falling after historical Proof of Play removal.","Logarithmic Y axis, rows read per hour. Observed hourly buckets (3 Sep ~11h–14h BRT) plus the 2 Sep overnight range. Not a fitted continuous series.","\u002Fimg\u002Fposts\u002Fdo-storage-storm\u002Fdurable-objects-recovery-cliff.png",[21,550],{"tone":551},"after",[25,553,555],{"id":554},"the-storm-radar","The storm radar",[30,557,558],{},"We already had Cost Sense, a production job running every 15 minutes to estimate Cloudflare spend and compare it with recent baselines.",[30,560,561],{},"Before this incident, it was primarily designed to catch slower account-level cost drift.",[30,563,564],{},"That was not enough.",[30,566,567],{},"This failure did not look like a classic application outage.",[30,569,570],{},"It looked like one quiet Durable Object repeatedly reading its own database.",[30,572,573],{},"A generic alert saying:",[30,575,576],{},[577,578,579],"em",{},"\"Cloudflare spend looks high\"",[30,581,582],{},"would have been directionally useful.",[30,584,585],{},"But it would not tell the operator what to do next.",[30,587,588],{},"So we added a TenantDO storm radar.",[30,590,591],{},"Its goal is simple:",[30,593,594],{},[42,595,596],{},"Detect an unbounded storage pattern early and identify the lever that can bound it.",[69,598],{":breakout":71,":height":599,":n":600,":width":601,"alt":602,"caption":603,"src":604,":hideOnMobile":71},"768","7","2048","TenantDO storm radar flow from storage metrics through detection and classification to Slack alert and mitigation.","Cost Sense turns TenantDO storage telemetry into a classification (read storm or write storm) and an explicit mitigation path.","\u002Fimg\u002Fposts\u002Fdo-storage-storm\u002Ftenantdo-storm-radar-detect-classify-mitigate.png",[606,607],"storm-radar-steps",{},[233,609,611],{"id":610},"what-it-watches","What it watches",[30,613,614],{},"Every production tick now samples the Lucimark TenantDO namespace for:",[110,616,617,620,623,626,629,632],{},[113,618,619],{},"SQLite rows read;",[113,621,622],{},"SQLite rows written;",[113,624,625],{},"Durable Object active time;",[113,627,628],{},"recent hourly storage buckets;",[113,630,631],{},"Worker and Durable Object request volume as supporting signals;",[113,633,634],{},"comparison against recent production baselines.",[30,636,637],{},"The previous sample is stored so the monitor can calculate deltas rather than relying only on day-to-date totals.",[30,639,640],{},"We also guard against missing samples after deployment. Treating an absent previous value as zero could manufacture a fake billion-row spike.",[233,642,644],{"id":643},"three-ways-a-storage-storm-can-trip","Three ways a storage storm can trip",[30,646,647],{},[42,648,649],{},"Pace",[30,651,652],{},"If day-to-date consumption is already far ahead of where historical baseline suggests it should be, raise a finding.",[30,654,655],{},"This catches slower burns early in the day.",[30,657,658],{},[42,659,660],{},"Rate",[30,662,663],{},"Extrapolate the latest short-window delta.",[30,665,666],{},"If the current slope would produce an abnormal full-day result, raise a finding even if the cumulative total still looks harmless.",[30,668,669],{},"This catches sudden accelerations.",[30,671,672],{},[42,673,674],{},"Absolute hourly floors",[30,676,677],{},"Relative comparisons are not enough for cliffs.",[30,679,680],{},"So the radar has absolute floors calibrated from this incident.",[30,682,683],{},[42,684,685],{},"Read storm",[30,687,688],{},"≥ 10 million billed rows read\u002Fhour, sustained across two consecutive hourly buckets.",[30,690,691],{},"Healthy TenantDO activity is normally around the low hundreds of thousands to roughly one million reads per hour.",[30,693,694],{},"The incident produced 50–70M\u002Fhour even during its quieter morning phase and as much as 0.3–1.4B\u002Fhour overnight.",[30,696,697],{},"The point of 10M is to alert far below the level we experienced.",[30,699,700],{},[42,701,702],{},"Write storm",[30,704,705],{},"≥ 2 million billed rows written\u002Fhour, also sustained.",[30,707,708],{},"Writes are the expensive meter.",[30,710,711],{},"A healthy Lucimark day produces only a fraction of that number in total. Millions of writes packed into individual hours therefore deserve immediate attention even before a relative baseline catches up.",[233,713,715],{"id":714},"detection-should-name-the-lever","Detection should name the lever",[30,717,718],{},"When the radar fires, the alert is built for someone who has not spent the previous twelve hours reading migration code.",[30,720,721],{},"It contains:",[110,723,724,727,730,733,736],{},[113,725,726],{},"current Durable Object storage metrics;",[113,728,729],{},"the abnormal finding;",[113,731,732],{},"directional cost information;",[113,734,735],{},"recent baseline;",[113,737,738],{},"the appropriate first-response control.",[30,740,741],{},"A read storm points first to the alarm\u002Fhousekeeping kill switch.",[30,743,744],{},"A write storm points first to the ingest\u002Fwrite-amplification controls.",[30,746,747],{},"This is the difference between observability and mitigation.",[30,749,750,753,754,757],{},[577,751,752],{},"\"Spend is high\""," is an observation.\n",[577,755,756],{},"\"Park the alarms\""," is a runbook.",[30,759,760],{},"On a loop that does not naturally stop itself, the second is what bounds the incident.",[233,762,764],{"id":763},"backstops","Backstops",[30,766,767],{},"A 15-minute production monitor is useful, but no single monitoring path should be trusted as the only line of defense.",[30,769,770],{},"We added two additional mechanisms.",[30,772,773],{},[42,774,775],{},"Daily overnight check",[30,777,778],{},"The daily operational snapshot independently checks for abnormal Durable Object read rates and can raise a separate notification.",[30,780,781],{},"The goal is to catch the class of problem most likely to get a long head start while nobody is watching.",[30,783,784],{},[42,785,786],{},"Backfill-state inventory",[30,788,789],{},"Namespace metrics tell us that the TenantDO fleet is abnormal.",[30,791,792],{},"They do not necessarily identify which tenant is responsible.",[30,794,795],{},"So the snapshot also inspects Proof of Play migration state across active tenants and reports incomplete or unknown migrations.",[30,797,798],{},"The states are surfaced worst-first:",[30,800,801],{},"indexing → deduping → updating → pending → done",[30,803,804],{},"Timeouts are classified as unknown rather than healthy.",[30,806,807],{},"That is the inventory we did not have on 2 September.",[30,809,810],{},"At the time, thousands of identical error lines had to be traced back to one tenant manually.",[233,812,814],{"id":813},"after-every-deploy","After every deploy",[30,816,817],{},"Production API, app, and player deployments now arm a temporary higher-sensitivity Cost Sense watch.",[30,819,820],{},"During that window:",[110,822,823,826,829,832],{},[113,824,825],{},"pace thresholds tighten;",[113,827,828],{},"rate thresholds tighten;",[113,830,831],{},"abnormal Durable Object behavior becomes easier to trip;",[113,833,834],{},"any alert includes deploy context.",[30,836,837],{},"This migration entered production as a routine API deploy.",[30,839,840],{},"A future storage regression should not receive an overnight head start simply because HTTP error rates look acceptable.",[25,842,844],{"id":843},"what-changed-in-our-engineering-rules","What changed in our engineering rules",[30,846,847],{},"What already holds on the platform, in brief:",[110,849,850,853,861,864,867],{},[113,851,852],{},"TenantDO boot stays cheap schema only — no large-table DML or expensive uniqueness rebuilds;",[113,854,855,856,860],{},"heavy backfill is asynchronous, batched, resumable, marked ",[857,858,859],"code",{},"done",", and safe to re-run after a reset;",[113,862,863],{},"cooldowns on O(n) states;",[113,865,866],{},"test the tail: the largest real dataset, not the median;",[113,868,869],{},"Cost Sense storm radar, alarm kill switch, post-deploy watch, and backfill inventory.",[871,872],"shipped-vs-open",{":shipped":873,":stillOpen":874},"[\"Cheap TenantDO boot policy\",\"Cooldown and helper index in the state machine\",\"Migration cadence decoupled from 30 s\",\"Storm radar (pace, rate, hourly floors)\",\"Alarm \u002F housekeeping kill switch\",\"Post-deploy Cost Sense window\",\"Overnight check and backfill inventory\"]","[\"Circuit breaker after repeated UNIQUE failures\",\"Incremental dedupe instead of full-table GROUP BY\",\"Staging soak with ≥500k rows on the PoP schema release gate\",\"Pause one tenant's migration without parking all housekeeping\",\"Rebuild-from-scratch when grain makes historical migration irrational\",\"Per-tenant SQLite cost attribution\",\"Safer tooling for destructive exits\"]",[30,876,877,878,881],{},"A circuit breaker here is a limit that ",[42,879,880],{},"stops"," retry after N failures — pause, alert, and require a state change — instead of letting the tick repeat forever. It is not yet on the UNIQUE path; it sits in the right-hand column.",[30,883,884],{},"The radar does not make a bad migration inexpensive. It reduces how long a bad migration can remain unbounded.",[25,886,888],{"id":887},"what-we-would-tell-another-team-running-sqlite-on-durable-objects","What we would tell another team running SQLite on Durable Objects",[30,890,891],{},"Most of that is already in the table. What does not fit a two-column grid:",[30,893,894,897],{},[42,895,896],{},"Split observability by failure surface."," HTTP charts will not tell you SQLite is scanning itself. Watch rows read \u002F rows written directly.",[30,899,900,903,904,907],{},[42,901,902],{},"Failed DDL is a state-machine transition."," If ",[857,905,906],{},"CREATE UNIQUE INDEX"," can fail inside a recurring worker, design that transition before shipping.",[30,909,910,913],{},[42,911,912],{},"A useful alert names the lever."," Not “spend is high.” Yes: “this pattern matches an alarm-driven scan loop — park the alarms.”",[30,915,916],{},"Hope is not a circuit breaker.",[25,918,920],{"id":919},"closing","Closing",[30,922,923],{},"We shipped a Proof of Play data model change whose migration strategy underestimated one large TenantDO.",[30,925,926],{},"The first failure mode was expensive writes.",[30,928,929],{},"The second was more revealing: a 30-second alarm had turned O(n) migration work into a loop with no natural stopping condition.",[30,931,932],{},"We stopped the multiplier, fixed the state machine, removed historical data that did not justify an open-ended rewrite, and restored the system to its normal range.",[30,934,935],{},"Then we changed the platform so a similar failure should be visible much earlier.",[30,937,938],{},"The most important lesson was not the number of rows.",[30,940,941],{},"It was the shape of the failure.",[30,943,944],{},"A recurring operation whose cost scales with table size needs a bound before it reaches production.",[30,946,947],{},"For Lucimark, that now means stricter migration mechanics, cooldowns, kill switches, realistic high-volume testing, and a radar that does more than report that cost is rising.",[30,949,950],{},"It tells us where the amplifier is — and which lever stops it.",[30,952,953],{},"Lucimark's job is to let operators run screen networks with confidence.",[30,955,956],{},"That includes the parts nobody sees on the wall: migrations, rollups, storage behavior, and the cost of being wrong about a loop.",{"title":958,"searchDepth":959,"depth":959,"links":960},"",2,[961,962,963,964,965,972,973,974,980,987,988,989],{"id":27,"depth":959,"text":28},{"id":89,"depth":959,"text":90},{"id":148,"depth":959,"text":149},{"id":201,"depth":959,"text":202},{"id":230,"depth":959,"text":231,"children":966},[967,969,970,971],{"id":235,"depth":968,"text":236},3,{"id":265,"depth":968,"text":266},{"id":293,"depth":968,"text":294},{"id":335,"depth":968,"text":336},{"id":372,"depth":959,"text":373},{"id":428,"depth":959,"text":429},{"id":441,"depth":959,"text":442,"children":975},[976,977,978,979],{"id":445,"depth":968,"text":446},{"id":464,"depth":968,"text":465},{"id":499,"depth":968,"text":500},{"id":512,"depth":968,"text":513},{"id":554,"depth":959,"text":555,"children":981},[982,983,984,985,986],{"id":610,"depth":968,"text":611},{"id":643,"depth":968,"text":644},{"id":714,"depth":968,"text":715},{"id":763,"depth":968,"text":764},{"id":813,"depth":968,"text":814},{"id":843,"depth":959,"text":844},{"id":887,"depth":959,"text":888},{"id":919,"depth":959,"text":920},{"src":991,"alt":992},"\u002Fimg\u002Fposts\u002Fmudanca-schema-alarm-30s-noite-bilhao-linhas\u002Fmudanca-schema-alarm-30s-noite-bilhao-linhas-cover.png","Illustration of a schema node, alarm pulses, and an amber data storm against a dark background.","How a Proof of Play migration on Cloudflare Durable Objects became an unbounded storage loop — and the radar we built so the next one cannot hide.",null,false,"md",[998,1002,1005],{"value":999,"label":1000,"hint":1001},"~625 million","billed rows written","SQLite writes, not money and not unique records",{"value":1003,"label":1004},"~30 s","alarm repeat",{"value":1006,"label":1007},"Screens playing","playback preserved",{},true,"\u002Fen\u002Fposts\u002Fmudanca-schema-alarm-30s-noite-bilhao-linhas","2026-09-04",{"title":14,"description":993},"schema-change-30-second-alarm-billion-row-night","en\u002Fposts\u002Fmudanca-schema-alarm-30s-noite-bilhao-linhas",[1016,1017,1018,1019],"cloudflare","durable-objects","proof-of-play","cost-observability","pt","schema-change-do-storage-storm","human","1OziTuffEZ6L6WkMIjWKSytGoRS5GWkVtdzSiUcnhmM",{"paulo-ramires":1025},{"handle":16,"name":1026,"role":1027,"avatar":1028,"description":1029},"Paulo R.","Founder and CTO","\u002Fimg\u002Fauthors\u002Fpaulo-ramires.jpg","Founder of Lucimark. Builds the platform end to end — from the Android player on the screen to the Workers serving the API.","mudanca-schema-alarm-30s-noite-bilhao-linhas",[1032,1045],{"path":1033,"title":1034,"description":1035,"slug":1036,"translationKey":1037,"category":8,"tags":1038,"authors":1040,"publishedAt":1041,"updatedAt":994,"draft":995,"featured":995,"cover":1042,"ogImage":1043,"discussionUrl":994,"translationStatus":1022,"readingMinutes":968},"\u002Fen\u002Fposts\u002Fplataforma-dooh-cloudflare-native","Why Lucimark runs entirely on Cloudflare","Workers, D1, R2 and one Durable Object per customer. What we gained, what hurt, and why there is no Docker anywhere in our stack.","dooh-platform-on-cloudflare","plataforma-dooh-cloudflare-native",[1016,1039,1017],"architecture",[16],"2026-08-12",{"src":1043,"alt":1044},"\u002Fimg\u002Fposts\u002Fplataforma-dooh-cloudflare-native\u002Fplataforma-dooh-cloudflare-native-cover.png","Illustration of a distributed screen network around abstract layers of compute, database, media, and state.",{"path":1046,"title":1047,"description":1048,"slug":1049,"translationKey":1050,"category":9,"tags":1051,"authors":1056,"publishedAt":1057,"updatedAt":994,"draft":995,"featured":1009,"cover":1058,"ogImage":1059,"discussionUrl":994,"translationStatus":1022,"readingMinutes":1062},"\u002Fen\u002Fposts\u002Frede-5-telas-receita-custos-payback-dooh","Can a 5-screen network make money? Revenue, costs and payback for a DOOH operation","How much can a small 5-screen DOOH network bill? See public market references, investment, costs, advertisers needed and payback scenarios.","5-screen-network-revenue-costs-payback-dooh","rede-5-telas-receita-custos-payback-dooh",[9,1052,1053,1054,1055],"dooh-network","indoor-media","payback","revenue",[16],"2026-09-05",{"src":1059,"alt":1060,"credit":1061},"\u002Fimg\u002Fposts\u002Frede-5-telas-receita-custos-payback-dooh\u002Fpraca-comercial-conectada-ao-amanhecer.jpg","Small digital media network connecting screens across different commercial venues","A small network can share infrastructure and inventory across different commercial points.",9,1788649408663]