Accelerating Transformer Training with NVIDIA Transformer Engine, Fused Kernels, BF16, FP8, and GPU Benchmarking

In this tutorial, we explore how NVIDIA Transformer Engine accelerates transformer workloads by combining fused GPU kernels, BF16 computation, and hardware-aware FP8 execution. We begin by installing Transformer Engine and detecting the active GPU architecture so that we can determine whether the runtime supports TE kernels, FP8 tensor cores, or only the pure-PyTorch fallback path….

Read Full News

AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs

AMD released Instella-MoE-16B-A3B, a fully open Mixture-of-Experts language model trained from scratch on Instinct MI300X and MI325X GPUs. The model holds 16B total parameters but activates only 2.8B per token. AMD is publishing weights from every training stage, along with data mixtures, training configs, and inference code. Two systems-level choices carry the release: Gated Multi-head…

Read Full News

Should you still buy your next smartphone — or subscribe to it instead?

The smartphone industry’s next battleground may not be the phone itself, but how consumers get it. As premium devices become more expensive, Apple, Samsung, and others are betting that leasing, subscriptions, and guaranteed buyback programs can make upgrading more attractive. This week, Apple launched Apple Upgrade in the U.S. in partnership with Klarna, allowing consumers…

Read Full News

What’s the best handheld mini fan?

It’s summer. You are sweaty. There’s some unprecedented heat wave that’s even worse than the last unprecedented heat wave. You are uncomfortable. You’re getting sweatier by the second. What do you do? This predicament — one that’s becoming more and more common — just flat out sucks. But if you have a handheld mini fan…

Read Full News