Skip to content
System One

Laya-MLX

An independent MLX port of Laya for Apple Silicon, 13.4ms median for a short English decision and 7.4ms on the multilingual checkpoint, with a Snake demo that plays at up to 75 moves per second under 1GB of memory.

Still from Laya-MLXPlay the video in the post below
Still from x.com

mizorewww/laya-mlx on GitHub

Stars
6,079
Forks
463
Language
Python
License
Apache-2.0
Last push
22 Sep 2026

Post on X

Views
4,195,937
Likes
14,024
Reposts
1,198

Read from GitHub and X on . Counts change daily.

Reported by the author

Short decision
13.4 ms P50One short question, P50 13.42 ms 7.39 msgithub.com
Snake demo
75.4 moves/s75.40 moves/s across 2,400 moves, zero deathsgithub.com

Laya-MLX ports Laya, an open typed-decision model, to Apple’s MLX framework so it runs natively on Apple Silicon with no PyTorch, no Transformers runtime and no cloud API call. On an M3 Max, the author measured a median of 13.42 milliseconds end to end for a single short English decision, and 7.39 milliseconds on the smaller multilingual checkpoint, with zero output tokens generated either way since the model reads option probabilities directly rather than writing an answer.

The Snake demo drives every move through a real Laya call rather than pre-scripting behavior, with a separate safety layer that can override an unsafe proposed move. With the tested compilation and prefix-reuse path enabled, the author recorded 75.40 moves per second across 2,400 moves, zero deaths, and two visible safety interventions in the paired test run, about 6.5% faster than the same run without that optimization. Peak memory for a single short decision stayed under one gigabyte.

The author’s X post, in Chinese, calls the port 50 times faster than Jev and says it runs on your own device, describes Laya as an open-source, Jev-like classifier built on text output probabilities, and says it plays Snake at 60 decisions a second. Install with pip from PyPI; a Hugging Face weights link and a Chinese-language readme are both linked from the repository.

How it uses Jev

    """
    decision = self.route(state, questions, model=model, task=task, lang=lang)
    agent = self.load(decision["model"])
    result = agent.system_one(state, questions)
    result["routing"] = dict(decision)
    return result

system_one = predict

def __repr__(self):
    return "Router(loaded=%s, max_loaded=%d, default=%r)" % (
        self.loaded,
        self.max_loaded,
        self.default,
    )

View in laya_mlx/router.py

Open on GitHub

More like this

Use cases this is tagged with