Blog

The hottest AI news & a perspective on internal AI

The big moves from Claude, ChatGPT, Grok, Copilot, Gemini and NVIDIA — with a perspective for Vietnamese enterprises on data sovereignty, PDPL compliance and private internal AI.

A bright meeting room with slatted blinds, a woman in an orange blazer standing at the right pointing at a whiteboard covered in hand-drawn distribution curves, a pie chart and sticky notes, while three colleagues sit around a table with laptops, notebooks and markers
Internal AI05 Aug 2026

How long does an in-house AI rollout take? The 8–10 week plan, who the client has to assign, and the six variables that decide whether you finish on time

An earlier piece answered what an on-premise LLM costs. This one answers the other half: how long it takes. The 8–10 week plan Namtech publishes, who the client has to assign at each phase, the compliance deadlines that are fixed by law, and the six variables that decide whether you land at the start or the end of that range.

Read more →
Head-on close-up of six blade server trays mounted side by side in one chassis, their fronts covered in rectangular ventilation slots, metal release levers, green and amber status LEDs and drive labels reading 146GB 15k and 300GB 15K, the whole frame washed in cold blue light
Internal AI03 Aug 2026

MCP drops sessions: the 28 July 2026 spec makes in-house agents replicable — and here is the list you must fix before upgrading

On 28 July 2026 the Model Context Protocol shipped spec 2026-07-28: the initialize handshake and the Mcp-Session-Id header are gone and the protocol is now stateless. In-house MCP servers can sit behind a plain load balancer — but this is a breaking release, and here is what to audit first.

Read more →
Aisle inside a data centre with two rows of black mesh server racks, blue and red network cables running inside, and a technician's monitor cart at the far end
Security31 Jul 2026

What exactly did you just download? Cisco opens a provenance database of nearly 900 open models — and a standard for what "derived from" means

On 30 July 2026 Cisco launched the AI Supply Chain Provenance Explorer — a free database of nearly 900 open models, built on the Apache-2.0 Model Provenance Kit and the Model Provenance Constitution. Why model cards are not enough, and what to check before loading an open model into an on-premise system.

Read more →
Three people in business suits meeting around a white table in a bright office, one typing on a laptop, with notebooks and a phone on the table — illustrating an enterprise AI investment decision
Data sovereignty29 Jul 2026

From pilot to production: in July 2026, enterprise AI money moved to two places — the bridge to production and the data border

July 2026: Cognizant launched an EMEA AI Unit (28 Jul) to bridge pilots to production, Atos launched Sovereign Cloud (23 Jul), Cognizant and Domyn moved LLMs on-premise (2 Jul). Gartner forecast cited: 5% to 50% of cloud AI workloads going sovereign by 2029.

Read more →
An engineer seated in front of six monitors showing system logs in a modern control room, illustrating a security operations centre
AI policy27 Jul 2026

The EU Cybersecurity & AI Action Plan (7 July 2026): when AI is both the weapon and the shield — how businesses should read it

On 7 July 2026 the European Commission presented its Action Plan on Cybersecurity and AI: model evaluation before market entry, a European access Blueprint, a secure testing platform. What businesses should take from it.

Read more →
A technician plugging network cables into a switch in a server rack, illustrating self-hosting an open-weights model on-premise
Open models · Data sovereignty24 Jul 2026

GLM-5.2: the world's strongest open-weights model now ships under MIT — and it changes the on-premise AI equation for businesses

Z.ai's GLM-5.2 leads the open-weights field on Artificial Analysis (score 51), ships MIT-licensed weights and a 1M-token context. Why it changes the on-premise AI and data-sovereignty equation for businesses.

Read more →
Abstract glowing red and white digital circuit render illustrating a language model's compute layer
Gemini · AI cost · Data sovereignty22/07/2026

Gemini 3.6 Flash: 17% fewer output tokens, but ML processing is still global-endpoint only — how Vietnamese businesses should read it

Gemini 3.6 Flash launched 21 Jul 2026: output falls to $7.50 per 1M tokens and uses 17% fewer tokens, but ML processing is still global-endpoint only.

Read more →
Dark screen split into terminal panes full of red error log lines, illustrating server log review after a security incident
WordPress · Security · RCE · Patching21/07/2026

wp2shell: WordPress core flaw exploited in the wild — a two-CVE chain, patch to 6.9.5 / 7.0.2 now

WordPress core is hit by the wp2shell chain (CVE-2026-63030 + CVE-2026-60137), now exploited in the wild. Patch to 6.9.5 or 7.0.2, then verify it landed.

Read more →
Laptop showing colorful source code, illustrating a developer building on-prem RAG
RAG · Vector database20/07/2026

Choosing a vector database for on-prem RAG in 2026: pgvector, Qdrant, Milvus, Weaviate, Chroma compared

Compare pgvector, Qdrant, Milvus, Weaviate and Chroma for on-prem RAG by license, version and storage, plus a multilingual embedding model for Vietnamese.

Read more →
Server room with racks and a monitor cart
On-premise AI18/07/2026

The real cost of running an LLM on-premise in 2026: VRAM, GPU and electricity — a TCO breakdown for businesses

How much does self-hosting an LLM really cost? A TCO breakdown: VRAM per model, 24/7 GPU electricity at Vietnam rates, and a comparison against cloud API prices.

Read more →
Glowing blue server racks in a data center
Moonshot AI17/07/2026

Moonshot launches Kimi K3: the world's largest open-weight model at 2.8 trillion parameters — closing in on Claude Fable 5 and GPT-5.6 Sol

Moonshot AI launches Kimi K3: 2.8T parameters, 1M-token context, prices far below closed models. Benchmarks, pricing tables and what it means for businesses.

Read more →
World map with a global network of connections
AI Analysis15/07/2026

Microsoft brings Copilot processing in-country for 15 nations — why "data sovereignty" is pushing enterprises toward internal AI

Microsoft is expanding in-country processing for Microsoft 365 Copilot to 15 nations. What it means for data sovereignty, how it maps to Vietnam's Cybersecurity Law + Decree 53/2022, and when internal/on-premise AI is the real answer.

Read more →
Rows of blade servers in an internal data center
AI Analysis13/07/2026

Sovereign LLM on-premise: why 2026 is the year enterprises "bring AI home"

On 1 July 2026 a sovereign LLM running entirely on a telecom's own servers was delivered. Why 2026 is the year enterprises pull AI behind the firewall.

Read more →
Padlock over data streams, illustrating AI chatbot data leaks
AI Security08/07/2026

2026 AI chatbot leaks: from McKinsey's Lilli to 300 million Chat & Ask AI messages — why on-premise AI is the way out

The two biggest AI chatbot data leaks of 2026 — 46.5 million internal McKinsey messages and 300 million from Chat & Ask AI — show why data must stay inside the organization.

Read more →
European Union flags waving against a blue sky
AI Regulation06/07/2026

EU delays high-risk AI rules to Dec 2027 (Digital Omnibus): why sovereign AI still matters

The EU has agreed to postpone obligations for high-risk AI systems (Annex III) from 2 Aug 2026 to 2 Dec 2027 via the Digital Omnibus (Parliament endorsed 16 June, Council cleared 29 June 2026).

Read more →
Light bursting at the end of a dark tunnel
Anthropic02/07/2026~6 min

US lifts the ban — Fable 5 is back globally from July 1: what did Anthropic commit to?

Commerce lifted export controls (June 30); Fable 5 returned globally July 1 across Claude.ai, Platform, Code and Cowork. In exchange: a >99% jailbreak classifier, a HackerOne program and deeper government coordination.

Read more →
8-step roadmap to build internal AI
Internal AI02/07/2026Series

Building your own internal AI: why & the roadmap (8 steps)

The map: what internal AI is, how it differs from cloud, and 8 steps from hardware to operations.

Read more →
Choosing hardware for internal AI
Internal AI02/07/2026Part 2/8

On-premise hardware for internal AI: how to choose

Apple Silicon vs GPU, and sizing by number of users (AI Box → Pro → Cluster).

Read more →
Choosing an open-source model
Internal AI02/07/2026Part 3/8

Choosing open-source models & commercial licenses

Qwen, Gemma, Llama, Mistral — what size, and why the license is make-or-break.

Read more →
Serving the internal AI model
Internal AI02/07/2026Part 4/8

Serving: install & optimize model speed

Ollama vs vLLM vs llama.cpp, quantization and an OpenAI-compatible API.

Read more →
RAG over internal documents
Internal AI02/07/2026Part 5/8

RAG: teach your internal AI your own documents

Embeddings, vector DBs and a pipeline so the AI answers with citations from internal docs.

Read more →
Integrating internal AI into workflows
Internal AI02/07/2026Part 6/8

Chat UI & integrating internal AI into your workflow

Open WebUI, OpenAI-compatible API, SSO/RBAC and real integration scenarios.

Read more →
Evaluating and tuning internal AI
Internal AI02/07/2026Part 7/8

Evaluating quality & tuning internal AI

Golden sets, reducing hallucination, safety guardrails and when to fine-tune.

Read more →
Operating and scaling internal AI
Internal AI02/07/2026Part 8/8

Operating, monitoring & scaling internal AI

Monitoring, backups, model updates and scaling AI Box → Pro → Cluster.

Read more →
Internal AI system architecture diagram
Internal AI02/07/2026

Internal AI system architecture, layer by layer

A full-system diagram + the data flow of one question, explained layer by layer on-premise.

Read more →
Internal AI security system
Security02/07/2026

The security system of internal AI: defense in depth

Network isolation, RBAC, encryption, audit logs, prompt-injection defense — layered protection.

Read more →
Trending Pool — periodic knowledge updates
Internal AI02/07/2026

Trending Pool: how internal AI stays current with world knowledge

How an isolated internal AI still refreshes world knowledge on a schedule via a controlled channel.

Read more →
How internal AI reasons
Internal AI02/07/2026

How does internal AI reason to answer accurately, like Claude?

Internal AI reasons the same way as Claude — next-token prediction, grounded on your docs via RAG. Explained simply.

Read more →
What is an AI token
AI basics02/07/2026

What is an AI token? Data units & the context window

What a token is, how the context window works, and why tokens drive AI cost.

Read more →
Departments & access control for internal AI
Internal AI02/07/2026

Departments & access control when rolling out internal AI

A department map + an RBAC matrix: each team sees only permitted docs, enforced at the RAG layer.

Read more →
Humanoid robot in a dark warm space
Anthropic02/07/2026

Anthropic launches Claude Sonnet 5: near-flagship, 1M context, $2/$10 promo

Anthropic's agentic mid-tier model (June 30): default for Free/Pro, 1M-token context by default, $2/$10 per Mtok promo until Aug 31.

Read more →
Locked iron gate in a dark corridor
OpenAI02/07/2026

GPT-5.6 launches like never before: the US government vets each customer

OpenAI unveils Sol/Terra/Luna (June 26) — initially only ~20 government-vetted partners can access Sol.

Read more →
Hooded hacker typing in the dark
Security02/07/2026

15 fake JetBrains plugins stole AI API keys from ~70,000 developers

Plugins posing as AI assistants on JetBrains Marketplace stole OpenAI/DeepSeek keys. All removed June 17.

Read more →
Close-up of a dark motherboard
OpenAI02/07/2026

OpenAI unveils its first AI chip "Jalapeño" with Broadcom: designed in 9 months

OpenAI's first self-designed inference ASIC (June 24) — 9-month tape-out, deploying late 2026. The hardware independence race.

Read more →
DSLR camera on a dark background
Google02/07/2026

Google launches Nano Banana 2 Lite: 4-second images at $0.034 per 1,000

Gemini 3.1 Flash-Lite Image (June 30): Google's fastest, cheapest image model for high-volume enterprises.

Read more →
Cybersecurity and system access
Anthropic22/06/2026

The US forces Anthropic to pull Fable 5 worldwide: the lesson for Vietnamese businesses

On 12/06/2026 the US government forced Anthropic to stop serving Fable 5 & Mythos 5; pulled globally within hours.

Read more →
Justice scales and a wooden gavel
AI Regulation23/06/2026

Decree 142/2026: Vietnam's first legal framework for AI

Decree 142/2026/NĐ-CP (effective 01/05/2026): 4-tier AI risk classification, conformity assessment, sandbox, a support fund. What businesses must prepare.

Read more →
Justice scales and gavel
Data Sovereignty23/06/2026

"The year of AI sovereignty": the EU AI Act tightens — why keep data on-prem

The EU AI Act tightens high-risk obligations from Aug 2026, fines up to 7% of global revenue. Why keep data on-premise.

Read more →
Servers in a data center
Grok23/06/2026

xAI brings Grok to Databricks: agents reason directly on your data

Grok available natively on Databricks (18/06) — agents reason over Lakehouse data without external pipelines.

Read more →
Professional camera collection
OpenAI23/06/2026

Getty Images partners with OpenAI: licensed images in ChatGPT

A multi-year display deal puts Getty images in ChatGPT (not training). Getty stock jumps.

Read more →
Microchip on a circuit board
NVIDIA23/06/2026

NVIDIA unveils the Vera Rubin platform: chips for the "AI factory" era

A new chip lineup spanning training to agentic inference — a base to build your own AI infra.

Read more →
Developer with multiple screens
Microsoft23/06/2026

Microsoft builds 7 in-house "MAI" models at Build 2026

Seven self-built models + a Copilot "super app" — even Microsoft is reducing reliance on OpenAI.

Read more →
Cybersecurity
Security23/06/2026

ServiceNow leaks data via a misconfigured endpoint

A config error exposed data beyond permissions — big SaaS still leaks. A lesson in data control.

Read more →
Phone showing ChatGPT
OpenAI23/06/2026

OpenAI retires GPT-5.2 for GPT-5.5, adds "Active sessions"

AI model lifecycles keep shrinking — a stability lesson for enterprises.

Read more →
A solid-state drive (SSD) on a dark surface
OpenAI23/06/2026

OpenAI Codex CLI bug silently wears down your SSD: ~640 TB/year

A Codex CLI logging bug writes ~37 TB in 21 days (≈640 TB/year) — enough to exhaust a 1TB SSD's endurance in under a year. A patch now cuts logging ~85%.

Read more →
Laptop showing programming code
OpenAI22/06/2026

ChatGPT gets a new "automated assistant": schedule reminders, monitor the web and apps

OpenAI launched Scheduled Tasks (17/06), with a dedicated "Scheduled" page, and retired Pulse.

Read more →
Server room with networking equipment
Grok22/06/2026

xAI's Grok 4.3 lands on Amazon Bedrock: 1M-token context, customizable reasoning

The first time xAI is an official model provider on AWS Bedrock, aimed at the enterprise.

Read more →
Team meeting in an office
Copilot22/06/2026

Microsoft unveils "Microsoft IQ" at Build 2026: a context layer for AI agents

Work IQ, Web IQ, Foundry IQ, Fabric IQ — GA across GitHub Copilot, Foundry and Copilot Studio.

Read more →
Artificial intelligence illustration
Gemini22/06/2026

Gemini 3.5 Flash at Google I/O 2026: a Flash model that beats last gen's Pro

Faster, cheaper, strong at coding and AI agents — Google's new "default" model.

Read more →
AI robot interacting with a digital interface
NVIDIA22/06/2026

NVIDIA unveils the RTX Spark superchip: running AI agents and huge LLMs right on your PC

Run LLMs up to 120 billion parameters on a personal machine, no cloud needed — "reinventing the PC".

Read more →
Data center
DeepSeek22/06/2026

DeepSeek launches V4: an open-source model that "closes the gap" with the world's best

The V4 preview (V4-Pro & V4-Flash), 1M-token context, selective-attention slashing costs.

Read more →
Server racks in a data center
Mistral22/06/2026

Mistral Small 4: reasoning, multimodality and coding folded into one Apache 2.0 model

MoE 119B/6B active, 256k context, up to 40% faster than Small 3 — runs on-premise.

Read more →
Network of glowing connections on a dark background
Meta22/06/2026

Goodbye open Llama: Meta launches Muse Spark — its first proprietary AI model

Meta Superintelligence Labs' in-house model, going closed (API private preview).

Read more →
Developer at work
Perplexity22/06/2026

Perplexity launches "Brain": an AI memory that learns overnight for the Computer agent

A context graph distilled into an "LLM wiki" that self-updates overnight — Perplexity says it makes the agent smarter.

Read more →