$10 in starter credits free when you create an account
Oct 09, 2026
Written by: Connor Brown

Hoonify and Together AI both run open-weight models behind an OpenAI-compatible API. Change your base URL, keep your code, pay per token. For a lot of workloads, they’re interchangeable. The decision gets easy once you ask a different question than most comparison pages ask: after you send a prompt, who holds it, who could read it, and where does the model actually run? Those three answers separate these platforms more than any benchmark does.
If you’re evaluating both, a lot of the checklist comes out even.
You can build a working product on either one. What separates them is ownership: of the prompt, of the weights, and of the machine the inference runs on.
| Hoonify | Together AI | |
|---|---|---|
| Zero data retention | On by default | Available, off by default |
| Default behavior if you change nothing | Prompts dropped after the request | Prompts and responses stored, “may use them for product improvements” |
| Cost of turning retention off | Nothing to turn off | Disables passthrough models |
| Covers data from before you enabled it | Not applicable | No, forward-only |
| Trains on your data | No | Not without explicit opt-in |
| Private networking / VPC | Yes | Yes, enterprise, incl. EU regions |
| On-premises, your own hardware | Yes | Not offered |
| Fully air-gapped, no egress | Yes | Not offered |
| GLM-5.2 cached input | $0.18 / 1M | $0.26 / 1M |
| Gemma-4-31B, per 1M tokens | $0.12 in / $0.38 out | $0.39 in / $0.97 out |
| Cheapest model in catalog | $0.12 in / $0.38 out | $0.05 in / $0.20 out (GPT-OSS-20B) |
| Free to start | $10 in credits, no card | No published credit amount |
| Model catalog size | Four production models | 57+ across text, vision, audio, video |
Prices and policy terms captured from together.ai/pricing, together.ai/privacy, docs.together.ai/docs/zero-data-retention, and hoonify.ai/catalog on September 30, 2026.
Start with what happens if you sign up, paste in your API key, and change no settings. That’s how most evaluations actually run, and it’s the configuration most pilots stay in long after the pilot ends.
On Hoonify, prompts and completions are dropped when the request completes. Zero data retention is the default, not a setting. There’s nothing to find, enable, or remember to turn on before your first real prompt.
Together AI has zero data retention too, and their documentation is admirably direct about how it works. Three sentences from their own documentation and privacy policy, as published in late September 2026, are worth reading closely.
The first: “ZDR is not enabled by default.” An organization admin has to turn it on in settings.
The second describes what happens until they do: “Together stores the prompts you send and the responses models return, and may use them for product improvements.”
The third is the one that matters most in a regulated shop. Their privacy policy states that ZDR “applies only from the moment you enable it and does not affect any data processed prior.” Retention is forward-only. Every prompt from the six-week evaluation you ran before anyone thought to check the privacy settings is still whatever it was.
There’s a fourth detail worth knowing before you flip the switch. Their docs note that enabling ZDR “turns off passthrough models automatically, because passthrough requires retained prompts.” So on Together, privacy has a price paid in model access. Turn retention off and part of the catalog goes with it.
None of that makes Together careless. They built the control, they documented it clearly, and a team that reads the docs and configures it properly gets a real zero-retention posture. The difference is what happens to the team that doesn’t. A default is a decision the vendor makes on behalf of every customer who never opens the settings page, and most customers never do.
For a regulated buyer, the practical question in a security review isn’t “does the vendor offer zero retention.” It’s “what was the setting on the day my engineer ran a live customer record through it to see if this thing worked.” On Hoonify that question has one answer regardless of who was paying attention.
“Private” is the most overloaded word in this category. Every vendor claims it and almost none of them mean the same thing. There are five rungs, and they are not interchangeable.
| Deployment model | What it actually means | Hoonify | Together AI |
|---|---|---|---|
| Shared serverless | Multi-tenant, vendor’s hardware, vendor’s network | Yes | Yes |
| Dedicated endpoint | Single-tenant, still vendor’s hardware and network | Yes | Yes |
| Private networking | VPC-based, region-pinned, hardware you rent | Yes | Yes, enterprise, incl. EU |
| Your own datacenter | Your hardware, your building, outbound permitted | Yes | Not offered |
| Air-gapped | Your hardware, no network egress of any kind | Yes | Not offered |
Together climbs three rungs, and their EU region support is a real option for teams with EU data residency requirements. Hoonify climbs all five, on one API, with the same code at every level.
Requirements arrive as rungs. A review asking for single tenancy is satisfied at rung two. One asking for data residency is satisfied at rung three. One that says the workload cannot leave a facility is only satisfied at rung four or five, and no amount of VPC configuration gets you there.
Hoonify runs on TurbOS, a compute orchestration platform built to deploy modeling, simulation, and HPC workloads for national laboratories, where disconnected environments are the norm rather than the exception. Air-gapped deployment isn’t a feature we bolted on for a deal. It’s the environment the platform grew up in.
That range matters because most teams don’t know which rung they’ll need. They find out in month nine, when a customer’s security review lands or a contract adds a residency clause. On Hoonify that’s a deployment change. On a platform that stops at rung three it’s a migration to a different vendor, in the middle of a deal.
Together has published a SOC 2 Type 2 report and documents their controls around encryption, audit logging, access management, and incident response. That’s a real security program.
What an air-gapped deployment does is change the shape of the question rather than the answer to it.
On any hosted model, including a private-networking tier, your vendor sits inside your trust boundary. Their controls, their staff, and their infrastructure partners become part of your assessment, and their audit cycle gets pulled into yours at every renewal. That’s as true of our hosted tier as of anyone’s. You are, in a real sense, inheriting someone else’s compliance posture.
Run the same software disconnected on your own hardware and the vendor leaves the boundary. The question becomes whether the software runs inside an enclave you already had accredited, under controls you already operate, staffed by people you already cleared. You aren’t inheriting a posture. You’re running code inside your own.
That’s why sovereignty and certification aren’t substitutes. A SOC 2 report tells you a vendor’s controls were tested by an auditor. It doesn’t tell you the data can be somewhere the vendor cannot reach, and for some requirements that second property is the only one that counts.
If your prompts can contain export-controlled technical data, there’s a question that no retention setting answers, on our hosted tier or anyone else’s. None of this is legal advice, and the right call for your program is a conversation with your export control counsel. But the shape of the problem is worth understanding before a security review raises it.
Encrypted data sitting in a third-party cloud is a familiar problem, and compliance teams have well-worn ways of handling it: the data is encrypted end to end and the provider never holds the keys. Cloud inference is different in one basic way.
A model can’t run on ciphertext. To generate a completion, the provider’s GPUs need your prompt in plaintext, in their memory, on their hardware. For the duration of the request, a third party is holding your data in the clear. Storage and processing look similar in a vendor’s marketing and are different in practice.
Zero data retention doesn’t change that, on either platform. Retention policy governs whether plaintext gets written down after the request. It says nothing about the fact that plaintext existed on third-party infrastructure during the request, or about who could have reached it while it was there. This is the limit of every retention promise in the category, ours included.
Air-gapped deployment removes the question instead of answering it. If the weights and the workload sit on your hardware, inside your facility, with no path out, there’s no third party in possession at any point. Whether that matters for your program is a question for your counsel. But it’s a very different conversation than walking an auditor through a cloud architecture.
One more dimension of ownership, and this one decides deals in a narrow set of fields.
Every hosted inference API sits behind the provider’s acceptable use policy and whatever safety tooling they run in front of the model. That’s appropriate for a shared service and it’s true of our hosted tier as much as anyone’s. It becomes a problem when legitimate work looks, from the outside, like the thing the policy is designed to catch.
A security team reproducing a published vulnerability, a patent firm prosecuting claims in a sensitive field, a defense analyst assessing an adversary’s capabilities: these are legitimate, funded jobs, and a general-purpose safety layer can’t always tell them apart from misuse. Teams in these fields have learned to expect refusals and hedged answers on exactly the queries that matter most to them.
Open weights running on your own hardware don’t have a provider-side policy layer, because there’s no provider in the path. The acceptable use policy that governs the work is the one your own compliance program writes, within the terms of the model’s license. For a cleared facility that is usually a tighter standard, not a looser one, because the rules are written for the work rather than for a general audience.
This is a real reason organizations move from a hosted API to on-premises deployment, and it has nothing to do with cost or latency. It’s about whether the tool is governed by rules written for the job.
Cost isn’t the argument here, but the numbers are worth stating plainly, including the one that goes against us.
| Model both platforms serve | Hoonify | Together AI | Difference |
|---|---|---|---|
| GLM-5.2 input | $1.40 | $1.40 | Identical |
| GLM-5.2 output | $4.40 | $4.40 | Identical |
| GLM-5.2 cached input | $0.18 | $0.26 | 31% less |
| Gemma-4-31B input | $0.12 | $0.39 | 69% less |
| Gemma-4-31B output | $0.38 | $0.97 | 61% less |
All rates per 1M tokens.
GLM-5.2 costs the same on both platforms. Our cached input runs cheaper, which helps on cache-heavy RAG workloads but isn’t a reason to switch by itself.
Gemma-4-31B is where the gap opens. On 100M input and 20M output tokens a month, that’s $58.40 on Together against $19.60 on Hoonify at published rates, about a third of the cost.
Together wins several categories outright.
Their catalog is far deeper. 57+ models spanning text, vision, image generation, text-to-speech, transcription, and video, against our four. If you need TTS and video from one vendor, or want to bake off a dozen models before committing, they have the inventory and we don’t yet.
They’re cheaper at the small end. GPT-OSS-20B at $0.05 / $0.20, Llama 3.2 3B at $0.06 both ways, DeepSeek V4 Flash at $0.14 / $0.28. Nothing in our catalog undercuts those. High-volume classification or routing on a small model, with data that isn’t sensitive, costs less on Together. That’s their win.
Their cluster offering goes further up the hardware stack, with GB200, GB300, and B300 at published rates plus preemptible pricing at a 50% discount for interruptible training. Self-serve fine-tuning starts at $0.48 per million tokens on a 16B LoRA without a sales call.
And their privacy documentation is genuinely good. The reason this post can be precise about their defaults is that they wrote them down plainly instead of burying them. A vendor who documents a tradeoff is easier to trust than one who markets around it.
If you need model variety, large training clusters, cheap small-model inference, or a certified multi-tenant cloud you can buy today with a card, Together is the better choice.
Choose Together if your data can live in someone else’s cloud. Their VPC and EU region options clear most security reviews, their catalog is deep, their small-model rates are the best in this comparison, and their zero-retention control works as documented once an admin enables it.
Choose Hoonify when ownership is the requirement rather than a preference. When retention has to be off on day one without anyone remembering to set it. When a contract, a regulator, or your own security team says inference happens inside your perimeter. When the technical data is export-controlled and you’d rather not have the cloud conversation at all. When the work is the kind a general-purpose safety layer gets wrong.
Together stops at rung three. We go to five, on the same API, with the same code.
Many of the teams that come to us started on a closed API, watched the bill grow, and then hit a compliance question their vendor couldn’t answer. Cost got them looking. Ownership is what made them switch.
You can test the cost and latency claims in about ten minutes. New accounts get $10 in credits, no card required, and zero retention is already on. Point your existing OpenAI client at Hoonify and compare. If the words air-gapped, ITAR, CMMC, or data localization appear in your requirements, start with a sizing conversation instead.
Start free · Talk to us about on-premises and air-gapped deployment
Hoonify brings mission-critical computing standards to AI. Built by founders from Cray, Sandia National Laboratories, and decades of high-performance computing experience, Hoonify comes from environments where the wrong answer is worse than no answer: national labs, scientific simulations, defense systems, energy research, life sciences, and other workloads that cannot fail quietly. Hoonify is built for organizations that need AI to be fast, private, dependable, and affordable at scale. With open weights, standard APIs, nothing retained, and infrastructure you control, customers keep ownership of their data, their hardware, and their roadmap.