95% utilisation doesn't mean 95% useful work. Cognit reads every layer of the compute stack — from silicon to scheduler — and tells you exactly which job is on which GPU, whether it's doing real work, and what it costs. All on-premise. No SaaS. No cloud dependency.
Cognit deploys inside your environment — on-premise, with no outbound telemetry — and connects to the tools you already run. No rip-and-replace. No data leaving your network.
No big-bang deployment. No data leaving your perimeter. We start with what you already have and show you value before asking for anything more.
From cluster health to compliance score to inference runtime — all visible, all governed, all on your hardware. Listed here are the modules backed by live telemetry on our own reference cluster. The product also ships procurement, tender and ESG-reporting surfaces; those are not listed, because they are not yet measured here.
Every Cognit deployment covers all three. These are the buyer entry points — where the pain is sharpest, the conversation starts.
Ten capabilities. One agent. Deployed in your environment — not ours.
Your GPU monitoring stack gives you device-level aggregates. It cannot tie kernel-level evidence back to the job or the tenant. That gap is the whole problem.
2·N·tokens/s ÷ peak FLOPS, against the dense BF16 peak. Against the 362 TFLOPS sparsity peak the same run reads 2.7% — we quote dense because vLLM cannot reach the sparsity figure, so that denominator flatters nobody honestly. AWQ is 4-bit, so this compares quantised throughput to a dense peak. Synthetic load, not a customer workload. 15 Jun 2026; raw scrapes retained./sys/class/hwmon. These measure different things by design; that is the point. Single node, idle. 10 Aug 2026.Three real workloads where the standard stack gives you a number. Cognit gives you the reason.
VK_ERROR_DEVICE_LOST — the renderer requires RT cores, and the H100 has none.
Bigger GPU ≠ faster.
Measured 21 Jun 2026 on real hardware. The agent already reads the card model and derives its RT-core capability, so it detects the mismatch — H100, A100 and V100 have no RT cores; L40S, A10 and RTX do. Gating placement on that verdict is in development — today the placement gate reads the declared job class.
Short versions only — each of these is measured, dated, and from a real machine: a customer's fleet or our own reference cluster. Ask for the long version on a call.
Renaming a process changes what it calls itself. It does not change what the kernel executed. Cognit identifies named scientific and AI workloads from kernel evidence — and is building the attestation layer that makes that identity something a tenant cannot forge.
exec -a changes what a process calls itself; it cannot change the inode the kernel ran. Full identity attestation across the library is in development.Wrapper-launched tools — Fluent, Abaqus, Parabricks — are recorded as exactly that: a wrapper legitimately shows an interpreter to the kernel, so we say identified, not attestable rather than pretending. The evidence ladder ranks every identity source by how hard it is to forge, and says so on the record.
Every finding is produced by symbolic rules over kernel evidence — the same inputs produce the same verdict, and verdicts carry the witnesses that produced them. A language model writes the plain-English explanation. It never decides the finding.
synthetic:false from a real source. The reference board hides by default anything unverified. An uploaded document can never move a verdict. The system is built to refuse a claim it cannot stand behind. Each of these is enforced in code and pinned by tests — a violation is a defect, and we say so.Forty frameworks in the control catalogue, assessed from live infrastructure evidence — not self-assessment checkboxes. Coverage depth is reported per control in the work paper, from the evidence present in your estate at the time of the run. Every gap is ranked by cross-framework impact with a clear remediation path.
synthetic:false from a real
source — a failed or empty read renders “unavailable”, never a green badge. And the board hides by default. A card is shown only if it names a real node and carries a real measurement, or is honestly empty — everything else is hidden rather than dressed up. On the default view of our own reference board, with that mode on, 15 of 39 cards render and 24 are hidden, including our own per-job trust score. Measured 10 Aug 2026. Panels currently modelled rather than measured: energy and inference.
A qualified auditor must review before reliance.
We don't ask for a platform commitment before you've seen your own numbers. Each step is fixed scope, fixed price, and credited in full against the next — so the only thing you risk is the first one.
Continuous observability and the compliance plane, deployed on your hardware — Observe for attribution and efficiency, Govern for the full compliance plane. Licensed per GPU, annually, and scoped to your cluster. Pricing on request — we quote it against what the audit found, not from a table.
Volume tiers apply above 64 GPUs and aggregate across clusters. Sovereign and defence deployments are scoped separately — talk to us.
Prices are set independently in each currency and hold for the term of the engagement — they are not converted at a daily rate. Revenue figures reported elsewhere on this page are stated in the currency actually received. Formal quotations are currently issued in INR; a quotation in your selected currency is available on request.
Most stacks bolt one tool per band and call it coverage. The runtime and device ends are well-observed. The kernel — where causation actually lives — is a blind spot.
★ ours — instrumentation we wrote and ship. ● kernel-direct — at least one signal at this layer is read from the kernel's own interface with no vendor daemon in the path; the evidence column names which, and what still comes through a vendor tool. ◆ integrated — read from the system of record, because Slurm is the authority on job state and asserting otherwise would make the evidence worse, not better.
Tools named are examples, not requirements. The four kernel-direct rows read identically whatever scheduler or GPU vendor you run — that is the point of reading the kernel.Nine layers read · one ours · four with a kernel-direct path · four from the source of record.
If you own the hardware, share GPUs across teams, and have ever wondered where the hours actually went — this is for you. Five situations, not five industries — and in every one of them, a gauge is lying to you.
Four objections from early conversations — paraphrased, not verbatim — and what each one taught us. We publish these because the people who say no teach you what the product is.
None of them is named here — partners get named when they say so, not before. Each engagement exists to prove one specific thing.
● live deployment · ○ in progress
Also in motion: an ongoing conversation with a major Indian GPU cloud · an early conversation with a global pharma manufacturer (GxP).
Cognit is an India-based infrastructure software company building the governance and observability layer for on-premise AI compute. We exist because the organisations that need this most — GPU clouds, enterprise AI platform teams, research and life-sciences computing — cannot send cluster telemetry to a foreign-jurisdiction SaaS dashboard, and no one was building for them.
We fund the deep-tech build with a working services practice — embedded systems, cloud communications and cybersecurity — which delivered and collected ₹13.3 lakh in FY 2024-25. Self-funded, debt-light and audited. We move carefully, we say what we mean, and we build in public where we can.
A working team, not a founder and a deck. Engineers who run the lab cluster, design the compute module and deliver on customer sites — and specialists we bring in for the parts that need a specialist: Rust systems and performance, hardware verification, audit, and robotics.
Share a few details and we'll get back to you within 2 business days. No demo-bot. No automated funnel. A real conversation with the team.
We'll be in touch within 2 business days at the email you provided.