Titan is coming. The GPU's reign is ending.
The limiting factor for frontier AI isn't compute — it's memory. Titan is an AI inference server answering with 8+ TB behind four Asimov chips built for transformer inference. Not graphics. Not training. Inference.
8 TB+
memory per system
11.8 TB/s
bandwidth
16 T
parameters per 4U box
10 M+
token context
An AI inference server built for memory, not FLOPs
Frontier models stopped being compute-bound some time ago. A modern LLM inference server spends most of its life moving weights and KV cache, not multiplying matrices — which is why GPU clusters sit underutilised while their memory subsystem saturates. Titan inverts the design: 8+ TB of memory behind 11.8 TB/s of bandwidth in a single 4U AI server, with four Asimov chips doing nothing but transformer inference.
Positron Titan specifications
- Memory
- 8+ TB per system
- Memory bandwidth
- 11.8 TB/s
- Accelerators
- 4× Asimov chips
- Model capacity
- Up to 16T parameters
- Context window
- 10M+ tokens
- Power
- 4.5 kW, air cooled
- Form factor
- 4U rack-mounted
- Parallelism
- Data, tensor, expert & pipeline
Titan — the AI inference hardware itself. Available 2027.
Asimov: the AI accelerator inside
- Memory per chip
- 864 GB – 2.3 TB LPDDR5x
- Bandwidth per chip
- 2.76 TB/s
- Power per chip
- ~400 W, air cooled
- Engine
- TransWarp, for transformer inference
Four of these per Titan. LPDDR5x gives roughly 6× the capacity of HBM at lower cost, without the power and cooling overhead of a traditional accelerator.
Frontier LLMs resident on one AI server
The interesting number is not peak throughput, it is what stays loaded. A frontier-class model in the two-to-three-trillion-parameter range occupies somewhere around 1–1.5 TB at four-bit weights. Against an 8 TB budget, that is a whole set of them resident at once on a single machine — each answering immediately, none paged in from disk or split across a fabric. Running the same set on a GPU alternative means a rack, a scheduler, and a conversation with your facilities team.
8+ TB
resident memory
The whole budget, in one address space rather than sharded across cards.
16 T
parameters supported
What that memory holds — roughly half a byte per parameter, so four-bit weights.
10 M+
token context
Enough headroom for whole codebases, long agent traces and document sets.
What that looks like in practice
These are the models a single 8 TB Titan is sized to hold simultaneously from 2027 — all resident together, not swapped in and out, and not spread across a cluster.
- Kimi K32.8T
- DeepSeek V4 Pro
- GLM 5.2
- Qwen3 Coder
- MiniMax
- Nemotron Omni
Why an inference accelerator beats a GPU rack
GPUs earned their position on training workloads, where raw FLOPs decide the outcome. Serving is a different problem with different economics, and it is the one most teams are now stuck paying for. This is where a purpose-built NVIDIA alternative changes the maths.
One box instead of a rack
Holding a multi-trillion-parameter model on GPUs means sharding it across many cards and often several nodes, then paying for that split on every token in interconnect latency. 8+ TB in a single chassis keeps the model resident in one address space — no cluster to schedule, no fabric to tune.
5× the tokens per dollar and per watt
Positron rates Asimov at 5× the tokens per dollar and per watt of NVIDIA Rubin. Inference is a margin business measured over years of continuous serving, so a multiple on tokens-per-watt compounds directly into unit economics.
Air cooled, so it fits your facility
The whole system draws 4.5 kW and is air cooled in 4U. No liquid loop, no CDU, no facility retrofit — the practical difference between deploying this year and negotiating with a landlord.
Built for inference only
Asimov’s TransWarp engine does transformer inference and nothing else. No graphics pipeline, no training-oriented silicon you pay for and never use. That narrowness is where the efficiency comes from.

Enterprise AI Infrastructure Partnership
Powered by Positron.ai - Purpose-Built Hardware for Next-Generation AI
Backed by a $230M Series B at a $1B+ valuation, Positron.ai builds purpose-built Transformer inference hardware. Our strategic partnership gives us access to cutting-edge, energy-efficient AI infrastructure that outperforms traditional GPU solutions.
Titan
Superintelligence in a Box - Coming 2027
4.5kW air cooled
Support for up to 16T Parameter Models with 10M+ token context window
8+ TB memory with 11.8 TB/s bandwidth
4x Asimov chips
Massive multi-node Data, Tensor, Expert, & Pipeline Parallelism
Asimov
Custom AI Accelerator Chip - Coming 2027
5x tokens/$ and tokens/watt vs NVIDIA Rubin
864GB to 2.3TB per chip (LPDDR5x)
2.76 TB/s memory bandwidth
TransWarp Engine for optimized Transformer inference
~400W air-cooled design
In the News
Recent Coverage & Milestones
Let's talk about how Positron hardware can accelerate your AI strategy
The honest details.
Forty-five minutes, technical, no deck. We go through the models you intend to serve and their memory footprint, your concurrency and latency targets, and where the machine would physically live. You leave with indicative pricing, a sizing view of how many systems your workload needs, and a position in the allocation queue. If Titan is the wrong answer for you, we will say so on the call.
Pricing depends on configuration, quantity and whether you are buying the hardware outright or having us host and operate it in the ANZ region. We quote after the sizing conversation rather than publishing a list price, because the number of systems you need depends entirely on the models and concurrency involved. Ask on the call and you will get a real figure, not a range.
Titan ships in 2027. Allocation is sequenced by when customers enter the queue, which is the honest reason to have the conversation now rather than in eighteen months. If you need capacity before then, we can serve you in the meantime on dedicated NVIDIA Blackwell GPUs in Auckland — see the near-term options below.
Your facility, a colocation provider of your choosing, or our Auckland facility hosted in Datacom Datacentres. Because Titan is air cooled at 4.5 kW in 4U, most existing racks can take it without a liquid-cooling retrofit — which is usually the deciding constraint on where AI hardware can go.
If we host it, the system sits in the ANZ region and your traffic stays there — no offshore failover, no cloud passthrough. If you own the machine, residency is whatever your own facility gives you. Either way inference runs on hardware dedicated to you, not shared multi-tenant capacity.
We do, as Positron’s APAC reseller — sizing, procurement, installation, and the serving stack on top. You deal with us in your timezone rather than filing tickets into another hemisphere, and we escalate to Positron directly where it needs silicon-level attention.
Reserve 2027 capacity
Titan ships in 2027 and allocation runs in the order customers join the queue. Booking a hardware call is how you enter it — and the call is worth having on its own terms, whether or not you end up buying.
Architecture review
What you are serving today, what it costs you, and where the memory ceiling actually bites.
Workload sizing
How many systems your concurrency and context targets need — and whether one is genuinely enough.
Indicative pricing
A real figure for the configuration you need, purchased outright or hosted and operated by us.
Allocation priority
Your place in the 2027 delivery queue. Sequenced by when you enter it, which is the whole point of doing this now.
Need inference capacity before 2027?
You do not have to wait for the silicon. We run dedicated NVIDIA Blackwell GPUs in Auckland today — your model weights, your private OpenAI-compatible endpoint, single-tenant and in-region. It is also how most customers bridge to Titan rather than sitting idle until it lands.