Search

38 results

Clear filters
  • SEPT. 30, 2026 / AI

    Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs

    To address the quadratic latency bottleneck of self-attention in high-resolution video diffusion models, developers implemented Sparse VideoGen (SVG) to dynamically route attention heads to highly structured spatial or temporal sparse masks. Translating this algorithmic sparsity into physical hardware speedups on TPUs required optimizing the Splash Attention kernel by bypassing empty memory tiles, restricting exact coordinate masking strictly to boundary tiles, and permuting token memory layouts into temporal-major order for contiguous access. By aligning these sparse masks with actual hardware tile execution, the combined optimizations significantly reduced wasted matrix operations and achieved up to a 1.69x end-to-end inference speedup for 1440p video generation.

    header (1)
  • SEPT. 24, 2026 / AI

    Reproducing Olmo 3 7B Pre-training in MaxText: case study of large scale training on TPUs

    The MaxText team successfully reproduced Ai2’s Olmo 3 7B language model from scratch on Google Cloud TPUs using JAX/XLA, precisely matching the original PyTorch-on-GPU reference across pre-training and mid-training stages on all held-out evaluations. The implementation achieved up to 57.4% Model Flops Utilization (MFU) and demonstrated robust infrastructure portability by surviving mid-run cluster resizes and cross-generation TPU shifts without requiring recipe alterations. Crucially, the exercise proved the necessity of comprehensive held-out validation by catching a silent data-loader memorization bug that artificially depressed training loss and would have otherwise faked a performance win.

    header
  • SEPT. 2, 2026 / AI

    4 engineering patterns behind the strongest AI Agents Challenge submissions

    The recent Google for Startups AI Agents Challenge revealed that the most successful multi-agent systems rely on foundational software engineering patterns rather than just raw model power. Winning architectures consistently implemented bidirectional MCP for seamless inter-agent communication, async event buses for parallel execution, strict unified validation for model fallbacks, and tiered routing to minimize expensive inference calls. By prioritizing these structural practices over simple linear prompt chains, developers can build more resilient, low-latency, and cost-effective agentic workflows.

    Gemini_Generated_Image_5wdk45wdk45wdk45
  • AUG. 27, 2026 / AI

    Decoding cosmic signals with deep learning and Keras

    Astroparticle physics sits at the exciting intersection of astrophysics and particle physics and stu...

    image_1
  • AUG. 13, 2026 / AI

    HeyGen x Google Cloud: Bringing Avatar IV to TPUs

    HeyGen ported their 18B+ parameter Avatar IV video generation model to Google Cloud's Trillium (v6e) TPUs via torchax and XLA, utilizing FSDP and Ulysses sequence parallelism across an eight-chip mesh. To achieve a 1.86x speedup for real-time streaming, the engineering team pipelined exposed all-to-all collectives, aligned sparse attention block sizes to eliminate mask padding, and bypassed softmax serial dependencies using a precomputed Cauchy-Schwarz upper bound. These custom Pallas kernel and compiler optimizations were deployed only after passing rigorous two-tier quality gates to guarantee byte-identical or mathematically equivalent pixel outputs.

    Gemini_Generated_Image_eshxr7eshxr7eshx
  • JULY 24, 2026 / AI

    Run Ray on TPU, Part 2: Ray AI libraries

    This second installment explores how Ray’s higher-level libraries—Serve, Data, and Train—abstract the complexities of running AI workloads on Google's TPU slices. Ray Serve uses a simple topology configuration to correctly gang-schedule large multi-host models, while Ray Data eliminates data-loading bottlenecks by feeding accelerators directly with native JAX batches. Finally, JaxTrainer streamlines distributed training across TPUs by automatically handling cross-slice coordination, checkpointing, and fault tolerance.

    header
  • JULY 21, 2026 / AI

    Scaling Agentic RL: High-Throughput Agentic Training with Tunix

    Tunix is Google’s new JAX-native post-training library designed to eliminate TPU idling bottlenecks when training multi-turn, tool-using LLM reasoning agents. It maximizes hardware throughput by combining highly concurrent, asynchronous rollouts with a decoupled producer-consumer pipeline, ensuring the trainer is constantly fed even while agents wait on network I/O or environment steps. Additionally, Tunix provides plug-and-play abstractions and continuous macro-level profiling, allowing developers to easily integrate custom open-source environments and optimize complex distributed workflows without massive code rewrites.

    Banner Image
  • JULY 8, 2026 / Mobile

    Bridging the Domain Gap: AI Race Coach built with Antigravity and Gemini

    On May 23, 2026, fresh off the stage at Google I/O, our Google Developer Experts (GDEs) converged on...

    GGBM3681 (1)
  • JUNE 22, 2026 / Web

    Measuring What Matters with Jules

    AI coding agents are rapidly shifting from reactive assistants that complete tasks when prompted to ...

    Measuring What Matters with Jules 1.0
  • JUNE 18, 2026 / AI

    How A2A is Building a World of Collaborative Agents

    Celebrating the first anniversary of the Agent-to-Agent (A2A) protocol, this blog post highlights how the framework enables autonomous AI agents to securely collaborate and hand off tasks without the rigidity of traditional APIs. By delegating complex workflows to specialized peer agents, A2A prevents context pollution, ensures data privacy, and simplifies application design through modularity. To demonstrate this ecosystem in action, the post spotlights FoldRun—an agentic interface for life sciences that orchestrates complex protein structure predictions—alongside diverse A2A use cases spanning commerce, data streaming, DevOps, and telecommunications.

    image2.original_6xqVyTd