Skip to content
Atlas Blog
LatestEngineeringBenchmarksDesign
atlasinference.io
blog.atlasinference.io

Notes from the inference layer

Kernel work, measured benchmarks, and what it takes to run frontier models on hardware you own. Everything we publish is reproducible from a commit.

All EngineeringBenchmarksReleasesDesign
Engineering Featured

DFLASH-2: the fastest single-machine numbers Atlas has produced

66.6 tokens per second on a stock build, one DGX Spark, one stream, and every figure reproducible from a commit.

Ronald R. Stesiak Sep 1, 2026 4 min read
  • Aug 31, 2026

    Seven Tenets Powering Atlas Inference Accelerated Workloads

    Atlas Inference is a free and open source LLM inference engine written from scratch in Rust. These are the seven philosophical tenets we started it on, and why we left the Python vLLM stack to do it.

    Engineering 7 min · TB
Atlas Inference Engine

Zero-trust inference on hardware you own. Pure Rust and CUDA, built in North Carolina.

Blog

  • Latest
  • Engineering
  • Benchmarks
  • RSS feed

Atlas

  • atlasinference.io
  • Documentation
  • Benchmarks
  • Download

Community

  • GitHub
  • Discord
  • X
© 2026 Atlas Inference · Community Edition AGPLv3
blog.atlasinference.io