Play the video in the post belowmizorewww/laya-mlx on GitHub
- Stars
- 6,079
- Forks
- 463
- Language
- Python
- License
- Apache-2.0
- Last push
- 22 Sep 2026
Post on X
- Views
- 4,195,937
- Likes
- 14,024
- Reposts
- 1,198
Read from GitHub and X on . Counts change daily.
Reported by the author
- Short decision
- 13.4 ms P50
One short question, P50 13.42 ms 7.39 ms
github.com - Snake demo
- 75.4 moves/s
75.40 moves/s across 2,400 moves, zero deaths
github.com
Laya-MLX ports Laya, an open typed-decision model, to Apple’s MLX framework so it runs natively on Apple Silicon with no PyTorch, no Transformers runtime and no cloud API call. On an M3 Max, the author measured a median of 13.42 milliseconds end to end for a single short English decision, and 7.39 milliseconds on the smaller multilingual checkpoint, with zero output tokens generated either way since the model reads option probabilities directly rather than writing an answer.
The Snake demo drives every move through a real Laya call rather than pre-scripting behavior, with a separate safety layer that can override an unsafe proposed move. With the tested compilation and prefix-reuse path enabled, the author recorded 75.40 moves per second across 2,400 moves, zero deaths, and two visible safety interventions in the paired test run, about 6.5% faster than the same run without that optimization. Peak memory for a single short decision stayed under one gigabyte.
The author’s X post, in Chinese, calls the port 50 times faster than Jev and says it runs on your own device, describes Laya as an open-source, Jev-like classifier built on text output probabilities, and says it plays Snake at 60 decisions a second. Install with pip from PyPI; a Hugging Face weights link and a Chinese-language readme are both linked from the repository.
How it uses Jev
"""
decision = self.route(state, questions, model=model, task=task, lang=lang)
agent = self.load(decision["model"])
result = agent.system_one(state, questions)
result["routing"] = dict(decision)
return result
system_one = predict
def __repr__(self):
return "Router(loaded=%s, max_loaded=%d, default=%r)" % (
self.loaded,
self.max_loaded,
self.default,
)

