SmallFireDragon Lab

Lab Articles

Deep dives into tech, design, and AI exploration

Publishing at 20:00 Daily, Yet Having Two "Todays": How Time Zone Boundaries in the Pipeline Create Duplicate Days
Article

Publishing at 20:00 Daily, Yet Having Two "Todays": How Time Zone Boundaries in the Pipeline Create Duplicate Days

This content pipeline publishes one article every day at 20:00 (SGT). The slug prefix is article-YYYYMMDD-. Before publishing, it checks the CMS: if a row with

Read More → →
A Batch Task Cost an Extra $200 in a Month: The Bill Exploded Before the Errors Did—Track Token Costs for Every Task
Article

A Batch Task Cost an Extra $200 in a Month: The Bill Exploded Before the Errors Did—Track Token Costs for Every Task

During last month’s reconciliation, we noticed that the daily cost of our batch export pipeline had risen from $0.30 to nearly $1.00. No one received any alerts

Read More → →
Pin Models and Prompts to Specific Versions: Don’t Let Upstream Silently Swap Engines and Capsize Your Tasks at Midnight
Article

Pin Models and Prompts to Specific Versions: Don’t Let Upstream Silently Swap Engines and Capsize Your Tasks at Midnight

At 3 AM, after an automated job had been running all night, the parsing success rate on the dashboard dropped from 99% to 71%. My first instinct was that someth

Read More → →
Task Stuck at 61% for Four Hours Undetected: Giving Long-Running Tasks a Heartbeat
Article

Task Stuck at 61% for Four Hours Undetected: Giving Long-Running Tasks a Heartbeat

Last week, one of our batch pipelines stalled. It started running at 3 AM. When I checked the dashboard at 11 AM, the progress was stuck at 61%. The task status

Read More → →
Don’t Let Models Parrot Your Secrets: Log Sanitization Needs Gates at Both Ends
Article

Don’t Let Models Parrot Your Secrets: Log Sanitization Needs Gates at Both Ends

Last month, the on-call bot I built for our inference service started gaining traction: when customers reported issues, it would pull relevant lines from produc

Read More → →
Three Nights of Autonomous Tasks, Doubled Bill: We Started Keeping a Ledger for Token Spending
Article

Three Nights of Autonomous Tasks, Doubled Bill: We Started Keeping a Ledger for Token Spending

At the beginning of this month, our autonomous delivery tasks ran continuously for three nights straight. Checking the bill in the morning, we found that the av

Read More → →
Three Lessons We Learned About Checkpointing from a Batch Job That Died at 98%
Article

Three Lessons We Learned About Checkpointing from a Batch Job That Died at 98%

Last month, we ran a batch job on our local inference server: classifying 24,000 anonymized customer service messages, with each message processed by a local mo

Read More → →
We Fell into the Circuit Breaker Trap Three Times
Article

We Fell into the Circuit Breaker Trap Three Times

Last week, we delivered an internal Q&A system for a client. The most embarrassing part wasn’t the model giving wrong answers, but the circuit breaker taking do

Read More → →
Running Evaluations in CI: A Model Upgrade Mishap and Its Fix
Article

Running Evaluations in CI: A Model Upgrade Mishap and Its Fix

Last week, we did something quite routine: we swapped the main model behind our local inference gateway from a 32B parameter model to a 70B one, and also bumped

Read More → →