Mix-FFN
Explore the latest content and insights.
Mix-FFN in SegFormer: Adding Local Context to Transformer MLPs
SegFormer Mix-FFN inserts a depthwise 3×3 convolution between two feed-forward projections, giving image tokens local spatial context without explicit positional embeddings.