On-Device AI
Explore the latest content and insights.
GPU có sẵn mà không dùng: bốn lỗi âm thầm khi huấn luyện trên máy người dùng
Huấn luyện đang dần quay về máy của người dùng, nơi phần cứng nằm ngoài tầm kiểm soát của nhà phát triển. Tôi đã để GPU trên máy Mac bị bỏ không suốt hai năm vì ứng dụng vẫn chạy được, chỉ là chậm. Khi bật Metal và đo đạc nghiêm túc, tôi phát hiện bốn lỗi hiệu năng không hề báo lỗi: is_bf16_supported() nhận nhầm khả năng phần cứng, gradient scaler âm thầm bỏ qua optimizer step, 16-bit chậm hơn 32-bit trên Apple Silicon, và giao diện báo dùng GPU trong khi mô hình thực tế chạy trên CPU.
The GPU Was Already There: Four Silent Bugs in On-Device Training
Training is moving back onto the machines people own, and those machines are not a fleet you control. Apple GPU support sat on my plan for two years because nothing looked broken; the runs finished, they were just slower. Turning Metal on was worth 1.7x to 4.2x per epoch on the one M1 I own, and much less end to end. The measuring is what found the real bugs, three of which never raised anything: is_bf16_supported() answering True on a card that only emulates it, a gradient scaler silently dropping optimiser steps and halving a mAP, and 16-bit being slower than 32-bit on Apple Silicon even though everything works.
Plan Once, Then Act: When the ReAct Loop Is the Wrong Harness for Small Local Models
On small local models, the standard ReAct loop has a failure mode nobody warns you about: the model calls one tool, declares victory, and stops. What we measured across 12 GGUF models in EdgeVox, why we added a plan-once dispatcher, and how to decide which loop your task actually needs.
Building EdgeVox: Chaining STT → Local LLM → TTS Without Touching the Cloud
A first-hand build narrative of EdgeVox, a fully offline voice agent that chains speech-to-text, a local LLM, and text-to-speech on one device. The architecture in plain language, ROS2 integration, the latency budget, and the failure modes nobody warns you about.
Câu chuyện chủ quyền AI của Việt Nam đang mắc kẹt ở một tầng quá cao 🇻🇳
Việt Nam đã có chip, có ba nỗ lực làm mô hình nền tảng tiếng Việt đang chạy, và bộ luật AI ràng buộc nhất Đông Nam Á. Vậy mà câu chuyện chủ quyền AI cứ xoay quanh một mô hình 70B. Khoảng trống thật nằm thấp hơn một tầng: đánh giá mở, dữ liệu sạch về giấy phép, mô hình chuyên biệt gắn với tuân thủ, và runtime chạy thẳng trên thiết bị. Từ 15/8/2026, Quyết định 33/2026/QĐ-TTg với 46 hệ thống AI rủi ro cao đã biến tầng đó thành nghĩa vụ pháp lý.
Vietnam's Sovereign AI Conversation Is Stuck One Layer Too High 🇻🇳
Vietnam already has the chips, three meaningful Vietnamese model attempts in flight, and the most binding AI law in Southeast Asia. The conversation about sovereign AI keeps demanding a 70B foundation model. The actual gap is one layer down: open evaluation, license-clean data, compliance-aware specialized models, and on-device runtimes that operationalize Law 134/2025 from March 2026.