Initialising · The Empyrean
Three Coders, One Held Back22 posts
← Back to InsightsAI Trends

Three Coders, One Held Back

August 15, 2026·7 min read
Three Coders, One Held Back · cover

Three coding models shipped this week. Two got cheaper. One stayed off the shelf.

Google released Gemini 3.7 Flash on 13 August at US$0.75/US$3.75 per million tokens — half the previous Flash-tier price, discounted through 31 December. DeepSeek's V4 Pro-0813 hit general availability the same day, adds a native OpenAI Responses API, and moves to peak/off-peak pricing on 17 August. On 14 August, Zhipu (operating as Z.ai) launched GLM-5.3 as the strongest open-weights coding model measured — but held the weights back two weeks after post-training produced cybersecurity capability the lab says outgrew its plan. At home, Bank Negara reported on 14 August that Q2 GDP grew 6.0% year on year — the strongest non-pandemic Q2 since 2014 — with headline inflation at 1.9%, OPR held at 2.75%, and nearly 80% of businesses reporting higher cost pressures.

Here's what happened, why it matters, and what your business should do about it.


The Big Three

1. Google Halved the Price of Its Coding Model — Again

On 13 August, Google released Gemini 3.7 Flash, the third Flash-tier update in about seven weeks. The stated targets are coding, agentic workflows and document processing; the unstated one is price. Through 31 December 2026, the API costs US$0.75 per million input tokens and US$3.75 per million output tokens; from 1 January 2027, both figures double. The introductory discount was also applied retroactively to Gemini 3.6 Flash — anyone who committed to 3.6 Flash three weeks earlier has been reset to half price on the newer version.

So what for a business owner? The rate at which mid-tier AI is getting cheaper has compounded three times inside a quarter — a straightforward argument against twelve-month AI service contracts at today's list price, and a stronger one against three-year deals. If your accounting suite, CRM, or customer-support platform embeds Gemini or another Flash-class model on your behalf, the price your vendor quotes is largely disconnected from the underlying model cost, and their invoice will lag the drop by six to nine months. Ask what model powers each AI feature you pay for, and when your vendor last renegotiated its rate with the lab.

2. DeepSeek's V4 Pro Went GA — With OpenAI's API Baked In

The same day Google shipped Flash, DeepSeek moved its V4 Pro model from preview to general availability as V4 Pro-0813. Two features matter. It now speaks the native OpenAI Responses API — meaning any tool built for ChatGPT's newer agent workflows can be redirected to a DeepSeek endpoint with minimal engineering. And from 17 August, DeepSeek introduces peak and off-peak API tariffs, mirroring how electricity utilities charge. The current standing rate is US$0.435/US$0.87 per million tokens; off-peak will run meaningfully lower.

So what for a business owner? A Chinese lab standardising on the OpenAI API means the switching cost between AI vendors is falling fast; the choice of underlying model is becoming a runtime decision, not an architectural one. The peak/off-peak shift points the same way: AI compute is being priced like a utility. Overnight batch jobs — invoice extraction, transcript processing, reconciliation runs — are about to become measurably cheaper than the same work triggered during office hours.

3. Zhipu Held Back Its Own Model's Weights After Training Went Further Than Planned

On 14 August, Zhipu AI (operating as Z.ai) released GLM-5.3, which it calls the strongest open-weights coding model measured to date. GLM-5.3 shares the same base as GLM-5.2 — its gains came from post-training alone, not fresh pretraining. Two details make the release unusual. The company reports that internal versions of the model helped security teams identify more than 2,400 vulnerabilities across 269 open-source projects, including flaws dating back decades. And Z.ai has held the weights back for roughly two weeks, saying cybersecurity capability emerged faster than expected and warrants a safety review before public release. The model is live via the company's API and coding plan; the open-weights promise, uncharacteristically for Zhipu, is on hold.

So what for a business owner? If a Chinese lab known for open publication is nervous about what its own model can do to under-patched software, that is a floor on how seriously you should take patch hygiene on anything facing the public internet. The commercial detail underneath: model capability can now jump substantially between releases from post-training alone — without the multi-hundred-million-dollar cost of a full pretraining run. Expect capability jumps to arrive more frequently and less predictably.

Closer to Home: Malaysia

The dated Malaysian story of the week is Bank Negara's 14 August release of Q2 economic data. GDP grew 6.0% year on year — the strongest non-pandemic second quarter since 2014 — on continued domestic demand and strong exports. Headline inflation edged up to 1.9%, still within BNM's 1.5–2.5% projection for the year, and the Overnight Policy Rate stays at 2.75%.

Behind the headline sits a less comfortable line from the same release: nearly 80% of surveyed businesses reported higher cost pressures in Q2. Roughly half said they may pass those costs on; the other half are absorbing. Fuel inflation ran at 5% in the quarter, reversing a -1.5% print in Q1, driven by the West Asia conflict feeding through import bills. The economy grew fast, borrowing stayed cheap, and margins tightened at the same time.

The compliance shape is unchanged. The Phase 4 grace period on RM1–5 million turnover businesses still runs to 31 December 2027, with the RM200–20,000 per-invoice penalty starting 1 January 2028. Transactions of RM10,000 or more already require an individual e-invoice today, and LHDN's February 2026 reporting flagged over 500,000 non-compliant cases and RM1.4 billion in unreported income — the grace period covers penalties, not audits. If you have not obtained the 12-character SSM BRN from every counterparty, do that this month.

Singapore, briefly: IMDA updated its Model AI Governance Framework on 12 August with new prompt-injection and jailbreak testing standards — relevant if you sell software into Singapore, otherwise inside-baseball.

What This Means for Your Business

  1. Do the margin math this quarter, not next year. BNM says four in five businesses like yours are seeing costs rise; half will pass them on, half will absorb. Pick your top ten SKUs or services, calculate the true landed cost today versus six months ago, and set a pricing decision date before September closes. The half that absorbs will report weaker margins by Q4; the half that raises prices without doing the math will lose volume.
  2. Assume your AI vendor's costs halved this week — and use that in your next contract talk. Google's newest coding tier just dropped fifty per cent; DeepSeek is about to charge off-peak rates. Whatever your accounting, CRM, or scheduling vendor charges for its AI features, the model cost underneath is falling faster than the invoice you receive. On your next renewal, ask two questions: what model powers this feature, and when did you last renegotiate with the lab. If the answer to the second is "more than six months ago," push for a shorter renewal with a mid-period price review.
  3. Patch the outside-facing systems this month. If an open-weights lab found 2,400+ vulnerabilities across 269 open-source projects using its own model, the same class of tools is already being pointed at your public e-commerce site, your customer portal, and your WhatsApp Business integration. Ask your web host and your e-commerce platform to confirm in writing that every software component is patched to the latest release, and that any WordPress or CMS plugins older than twelve months are updated or removed. That single email closes most of the easy exploit paths.

The Practical Question

If the price of AI is halving every quarter and my software vendor is holding me to a twelve-month contract at last year's rate, what is the actual cost of that lock-in?

The math is not complicated. If you pay RM800 a month for an AI-embedded tool and the underlying model cost has fallen roughly 50% since your renewal, a meaningful slice of that charge is margin your vendor is capturing because your contract has not caught up. Multiply by every AI-tagged subscription on your books, then by twelve months. Put a calendar note sixty days before every AI-related renewal to ask what has changed in the vendor's upstream contract. The ones that answer honestly are worth keeping; the ones that refuse are telling you where to look for savings.


At The Empyrean, we help Malaysian SMEs find the practical, repeatable tasks where AI delivers value without disruption. If you're not sure where to start, we're happy to take a look at your operations and tell you honestly what would make sense.

Talk to us →