Advanced Versatile Inference Silicon

ZilionX™ Transformer Engine

Developing an inference accelerator for Transformer workloads on constrained edge systems. The current RTL integrates attention, feed-forward and normalization paths. The Stage 3M memory fabric has been integrated with the Transformer RTL in Stage 3N. Streaming DMA and end-to-end simulation measurements are available; timing optimization and FPGA evaluation are next.

Current status · October 2026
Integrated RTL and an eight-layer simulation benchmark are available. Stage 3M memory fabric implementation is complete. Stage 3N memory-fabric RTL integration is complete. The checkpointed four-token, D=8, eight-layer test records 31,553 inference cycles and 33,892 load-to-result cycles including cache warm-up. Seven-suite regression and loader negative tests passed; the earlier results remain available on Benchmarks.

ZilionX development

Two foundations for the next stage

The hardware data path and the software path are advancing together. Each milestone below states what has been built and what is still being evaluated.

RTL integration complete · synthesis next

Memory fabric

The Stage 3M memory fabric is implemented and integrated with the Transformer RTL. The verified four-token, D=8, eight-layer checkpoint records 31,553 inference cycles and 33,892 cycles including DMA loading and cache warm-up. Regression and negative loader tests passed.

Follow integration progress
SDK v0.1 · model-package foundation

Developer SDK

The first SDK work defines and validates a ZilionX model package so model data can enter a controlled evaluation flow. Runtime integration, representative model execution and FPGA evaluation are subsequent milestones.

Discuss an evaluation

Engineering focus

Transformer datapath

Fixed-point attention and softmax RTL alongside integrated feed-forward, normalization and residual paths.

Data movement

Stage 3M memory fabric integrated with the Transformer RTL. DMA loading, cache warm-up and end-to-end cycles are recorded. A larger-token stress test passed with seven stall client-cycles; timing optimization and larger dimensions are next.

Evaluation path

Cycle-level simulation today; FPGA evaluation and representative model workloads are future milestones.

See the evidence

Architecture overview explains the blocks and intended memory integration. Benchmarks provides the measured cycle counts and test scope. Progress separates completed work from active and planned work.