tools.ml.box

Calculators for people who actually run the models

Small, fast, dependency-free tools for local LLM work. Real architecture configs rather than rules of thumb, so grouped-query attention, sliding windows and MoE routing all land where they should. Every result is a URL you can paste at someone.

Tools

    How these work

    Everything runs in your browser. There is no backend, no account, no analytics and no network request after the page loads — which is also why the whole configuration lives in the query string. That is the share button: the URL is the state.

    The numbers are estimates built from first principles, with constants calibrated against published llama.cpp and vLLM benchmarks. They are good enough to decide what to download and which card to buy, and they are not a substitute for benchmarking your own stack. Where a figure is shakier than it looks, the tool says so.

    Planned

    A public backlog, not a promise. Nudge me if one of these would actually be useful to you — that moves it up the list.