It runs on stock vLLM behind a small shim that serves Jev-format requests, and each answer takes one forward pass and one output token. H2O's recorded demos show it checking claims against a written policy, guiding browser agents and choosing moves in a Freedoom level, where a thin controller does the aiming.
How you can use it
H2O-Lightning-4B suits large piles of routine checks run on your own computer, such as whether a claim follows a written policy or what total a photographed receipt shows. Each answer comes with how sure the model is, and H2O lists how many answers would be passed to a person at different levels of certainty.