AI
Coming soonAI infrastructure without the complexity
Endpoints, inference and retrieval for developers. Private models and company knowledge for organisations. One key, one API, one invoice.
Tokens (30d)
1.4B
Cost (30d)
$527.22
Token usage by workload
- Chat812M · 57%
- Embeddings486M · 34%
- Vision124M · 9%
Endpoints
- Running64.2K · 780 ms
chat-production
Llama 3.3 70B
- Running1.1M · 42 ms
search-embed
Embed Large
Latency distribution
Developers
AI infrastructure for developers
Hosted endpoints, retrieval and model APIs behind one credential, so swapping a model is a configuration change rather than a rewrite.
AI Endpoints
Coming soonA stable, OpenAI-compatible URL per project. Change the model behind it without touching your code.
Inference
Coming soonServerless serving for open-weight models, designed for per-token billing with no idle cost.
Embeddings
Coming soonHigh-throughput embedding generation for search, clustering and recommendation pipelines.
Vector Databases
Coming soonManaged vector collections on the same private network as your endpoints, so retrieval never leaves Hostacker.
AI APIs
Coming soonTranscription, reranking, vision and moderation behind the same key and the same invoice.
AI products are not open for use yet. Pricing is published once inference capacity is contracted.
Enterprise
Private AI for enterprise
For organisations that want the capability without sending internal data to a shared, multi-tenant service.
Private LLM
PreviewIsolated language-model infrastructure intended for company data and internal applications.
Enterprise RAG
PreviewConnect company documents, procedures and knowledge to a private AI experience with access control.
Dedicated Inference
PreviewReserved or dedicated GPU infrastructure for AI workloads that need predictable throughput.
Also part of Enterprise AI
Enterprise knowledge indexing, MCP access for approved agents, and the security and governance controls that make both reasonable inside a company.
- Enterprise Knowledge
- Vector Infrastructure
- MCP
- SSO
- Audit logs
- Deployment policies
Platform
What the AI layer is built to give you
Not telemetry — Hostacker serves no traffic yet. These are the properties the platform is being designed around.
One key across every model
Endpoints, embeddings and model APIs are designed to authenticate the same way, so swapping a model is a configuration change.
Usage visible per project
Token consumption, request volume and cost are attributed per endpoint and per project rather than arriving as one opaque bill.
Retrieval on a private network
Vector collections are planned to sit beside your endpoints, so documents and embeddings never traverse the public internet.
Cost estimated before you commit
Every provisioning action is designed to return an estimate first, including the ones an AI agent initiates through MCP.
Models
Swap models without rewriting your integration
Endpoints are designed as stable URLs. Point one at a different model and your application keeps working.
| Model | Status |
|---|---|
Llama 3.3 70B Instructhk-llama-3-70b | Preview |
Llama 3.1 8B Instructhk-llama-3-8b | Preview |
Mistral Smallhk-mistral-small | Coming soon |
Qwen Coder 32Bhk-qwen-coder-32b | Coming soon |
Hostacker Embed Largehk-embed-large | Coming soon |
Ship AI features, not AI plumbing
Tell us what you are building and we will map it to the AI infrastructure it needs.