INSNAPSYS
AI infrastructure

Fullmetal

A distributed network for running open-source LLMs, without renting a datacentre.

Client
Fullmetal.Ai
Sector
AI infrastructure
Our role
Platform and distributed agent network build
Status
Live, released as a research preview

The brief

What was breaking

New open-source language models arrive weekly, and running them yourself is the only way to keep your prompts private. That is also the expensive way. Fullmetal set out to make model choice and privacy affordable for people building personal projects, by pooling machines instead of buying them.

  • The number of open-source models released each week makes choosing and testing one a project in itself.
  • Self-hosting an LLM requires hardware that is idle most of the time, which is difficult to justify for a personal project.
  • Existing AI APIs mean sending your prompts to a third party, which is exactly what people self-host to avoid.

What we built

The system we delivered

A distributed agent network behind one API

Requests are served by a pool of Fullmetal-hosted and community-contributed machines, each carrying at least 8GB of RAM or VRAM, so model availability comes from the network rather than from one operator's hardware budget.

Idle capacity that earns

Agent owners choose whether their machine serves public prompts while idle, and earn tokens for the prompts it handles.

Encrypted prompts and responses

Traffic is encrypted end to end so the privacy argument for self-hosting still holds when the compute belongs to someone else.

Agent operations in the open

Operators see each agent's status, creation date, prompts served, coins earned, and public-serving toggle, so participation in the network is legible rather than a black box.

In-browser inference with WebLLM

Models including Llama 2 and RedPajama-INCITE run as quantised builds directly in the browser, with cached parameters making subsequent visits fast.

Built with

The stack behind it

The lead tier is what makes this system what it is. Everything under it is the platform that carries it.

AI and models

WebLLMLlama 2 7BRedPajama-INCITEWizard-Vicuna4-bit quantisationDistributed inference

Platform

Agent orchestration APIAPI key managementToken accounting

Security

Encrypted prompts and responsesInvite codes

What changed

The result in production

Model choice without the hosting bill

Pooled community machines put a range of open-source models within reach of individual projects.

Privacy that survives the move off your own hardware

Encrypting prompts and responses removes the trade-off between not self-hosting and not sending your data to a third party.

An incentive to keep the network supplied

Token rewards for serving public prompts give idle high-spec machines a reason to stay in the pool.

The screens

What it looks like in use

Agent status: models served, prompts handled, coins earned, and the public-serving toggle per agent.
The chat demo, running a selected open-source model and returning generated code.
Hosting a model in-browser through WebLLM, with quantised parameters cached on first load.

Same problem?

Let's scope what this would look like for you

Start with a two-week Discovery Sprint. We map your highest-value workflows and deliver a prioritised pilot roadmap grounded in what we have already shipped.