Bo Li
AboutTech BlogsPublications
  • Tags
  • LLM-inference
Posted 2026-07-13Updated 2026-07-13Tech Blogs

Accelerating Long-Context Inference with Skip Softmax Attention

Skip Softmax Attention is a drop-in post-training sparse attention method and the productized implementation of BLASST in NVIDIA TensorRT-LLM. It is available for accelerating both LLM workloads and visual-generation workloads.

The original article was published in NVIDIA TensorRT-LLM and is embedded here for convenience.

Read more
Bo Li

Bo Li

DevTech Compute Engineer at NVIDIA

Posts

2

Category

1

Tags

7

Recents

2026-07-13

Accelerating Long-Context Inference with Skip Softmax Attention

Tech Blogs

2026-07-13

Optimizing MoE Communication with One-Sided AlltoAll Over NVLink

Tech Blogs

Categories

  • Tech Blogs2

Tags

BLASST1
GPU-communication1
LLM-inference1
MoE1
NVLink1
TensorRT-LLM2
sparse-attention1

Archives

  • July 20262
Bo Li

© 2026 Bo Li  Powered by Hexo & Icarus

Copyright (c) 2026 Bo Li

×