PyTorch

Explore the latest content and insights.

  • GPU có sẵn mà không dùng: bốn lỗi âm thầm khi huấn luyện trên máy người dùng

    GPU có sẵn mà không dùng: bốn lỗi âm thầm khi huấn luyện trên máy người dùng

    Huấn luyện đang dần quay về máy của người dùng, nơi phần cứng nằm ngoài tầm kiểm soát của nhà phát triển. Tôi đã để GPU trên máy Mac bị bỏ không suốt hai năm vì ứng dụng vẫn chạy được, chỉ là chậm. Khi bật Metal và đo đạc nghiêm túc, tôi phát hiện bốn lỗi hiệu năng không hề báo lỗi: is_bf16_supported() nhận nhầm khả năng phần cứng, gradient scaler âm thầm bỏ qua optimizer step, 16-bit chậm hơn 32-bit trên Apple Silicon, và giao diện báo dùng GPU trong khi mô hình thực tế chạy trên CPU.

  • The GPU Was Already There: Four Silent Bugs in On-Device Training

    The GPU Was Already There: Four Silent Bugs in On-Device Training

    Training is moving back onto the machines people own, and those machines are not a fleet you control. Apple GPU support sat on my plan for two years because nothing looked broken; the runs finished, they were just slower. Turning Metal on was worth 1.7x to 4.2x per epoch on the one M1 I own, and much less end to end. The measuring is what found the real bugs, three of which never raised anything: is_bf16_supported() answering True on a card that only emulates it, a gradient scaler silently dropping optimiser steps and halving a mAP, and 16-bit being slower than 32-bit on Apple Silicon even though everything works.