Skip to content
System One

kikoncuo/jevfire

Jev-inspired parallel decisions for CUDA LLMs: independent typed fields are batched through vLLM over one shared context prefix, single-token labels are scored with the pretrained model's own head, and JSON is assembled in code. No retraining. Includes a browser Super Mario run at 71 ms mean inference per action.

Open on GitHub

More like this

Use cases this is tagged with