Kernel-level · On-premise · Zero cloud calls

Your GPU dashboard
is lying to you.

95% utilisation doesn't mean 95% useful work. Cognit reads every layer of the compute stack — from silicon to scheduler — and tells you exactly which job is on which GPU, whether it's doing real work, and what it costs. All on-premise. No SaaS. No cloud dependency.

Cognit is a GPU intelligence layer that sits on your existing cluster — reads every signal from kernel to BMC — and tells you, in plain English, why jobs run slow, which GPU-hours are wasted, and what that costs. On-premise. No cloud. No rip-and-replace.
India-built No outbound telemetry Validated on NVIDIA
The 9-layer compute stack
Job / Workload
scheduler accounting
Observe
Framework
training profiler
Observe
Runtime / Serving
serving metrics · partial
Observe
Container
/sys/fs/cgroup
Direct
Scheduler
scheduler queue
Observe
Kernel (Linux)
eBPF agent
Cognit's layer
OS Userspace
chronyc · timedatectl
Direct
Firmware / BMC
/sys/class/hwmon
Direct
Hardware
/sys/bus/pci
Direct
Products

One control plane.
Every layer of your cluster.

Cognit deploys inside your environment — on-premise, with no outbound telemetry — and connects to the tools you already run. No rip-and-replace. No data leaving your network.

Commercial · Edition 1
Observe
Visibility + efficiency
Agentless — on-premise
Reads your existing Prometheus / DCGM. No agent to install. Connects in minutes. Kernel-band evidence needs the full agent.
Per GPU · on request
On-premise — full agent
All 9 telemetry bands including kernel eBPF. Zero data egress. No outbound telemetry.
Per GPU · on request
What's included
  • Every telemetry band fused — fd→GPU attribution, per-tenant usage
  • MFU vs util · io-wait (idle vs storage-starved)
  • DCGM / vLLM / scheduler correlation
  • Dashboards, alerts, and pattern detection library
  • GPU routing intelligence — NVIDIA today · vendor-agnostic by design
Connects to your existing stack —
Slurm
Kubernetes
Prometheus
DCGM
vLLM
Triton
Redfish / IPMI
Lustre / NFS
InfiniBand
Terraform
Ansible
Pulumi
AMD MIGraphX
OpenVINO
Warewulf / xCAT
How it Works

From first conversation
to full visibility — in days.

No big-bang deployment. No data leaving your perimeter. We start with what you already have and show you value before asking for anything more.

01
Starting point — interchangeable with 02
Share your cluster metrics
Tell us what you're running — GPU count, scheduler, utilisation numbers, job queue length, anything you have. We analyse it and tell you exactly where the waste or risk likely is before a single line of code touches your environment.
→ Export from Prometheus, DCGM, or Slurm sacct
→ Even a screenshot is enough to start
→ We identify the pattern, you confirm the pain
02
Starting point — interchangeable with 01
See it live on our cluster
Not slides. Not a recorded demo. A live walkthrough on our own lab node — live BMC and GPU telemetry read off real hardware, against the real control catalogue. You see exactly what the platform looks like on real hardware. Energy and inference panels are modelled today, and labelled as such.
→ 30-minute technical walkthrough
→ Kernel attribution, cost breakdown, compliance score — live, with modelled panels labelled
→ Bring your MLOps, infra, or compliance team
03
The proof point
Observability on your cluster — read-only
Cognit connects to your existing stack — Prometheus, DCGM, Slurm — in read-only mode. No agents on compute nodes to start. No data leaves your cluster at any point. Your team sees their own jobs, their own GPUs, their own cost and compliance posture — for the first time, at the kernel level.
→ Connects to your existing Prometheus / DCGM in minutes
→ Zero writes to your cluster · zero data egress
→ Written findings from your own exported metrics — before anything is installed
→ Full on-premise agent deployment only if you choose to proceed
Steps 01 and 02 can happen in either order — or simultaneously if you prefer to see the demo while we review your metrics.
What's inside

Every module. One screen.

From cluster health to compliance score to inference runtime — all visible, all governed, all on your hardware. Listed here are the modules backed by live telemetry on our own reference cluster. The product also ships procurement, tender and ESG-reporting surfaces; those are not listed, because they are not yet measured here.

Applies to the sixteen module descriptions below.
Workloads
Dashboard
Cluster health
GPU utilisation, active nodes, PCIe throughput, ECC errors, active patterns, node status, and recommended actions — in one executive + ops view. Answers three questions: where am I wasting GPU time, am I compliant, are my models healthy.Where am I wasting GPU time, am I compliant, and are my models healthy? One screen, three answers.
Patterns
Detection library
Pre-built detection patterns across scheduling, memory, thermal, hardware, storage, network, jobs, and accounting. Ships ready — no configuration. Each pattern is status-tracked: ACTIVE / WARNING / CLEAN / PENDING — a pattern that cannot run on your hardware reports PENDING and says why, rather than quietly returning CLEAN.What's going wrong that nobody has noticed yet? Known ways clusters go wrong, watched from day one — nothing to configure.
Jobs
Kernel attribution
Role-scoped job view with Kernel Attribution — which job is driving or starving each GPU at the kernel level. Pramaan Trust Score per job — per-workload scoring at admission. In development.Which job is hogging the GPU, and which one is starving? The operating system knows. Now you do too. Per-job trust scoring at admission is in development.
Job Analytics
GPU heatmap · MFU
GPU utilisation heatmap by partition and hour. Partition efficiency summary with MFU% — the real model-FLOP utilisation number your CFO wants.Which hours and which machines actually earned their keep? A picture your finance team can read without a translator.
Scheduler
Herd-intelligent
Herd-Intelligent Job Placement — nodes scored via live ECC history, PCIe health, and NCCL herd patterns before job submission. Flags degraded hardware; the flag is advisory and the operator decides. Automatic prevention is in development.Is this machine healthy enough to run the job? Checked before the job starts, not after it fails. Advisory today; enforcement is in development.
GPU Routing
validated on NVIDIA · AMD next
Workload-to-GPU routing recommendations by job class — LLM inference to high-bandwidth memory, CUDA training to CUDA-optimised silicon, latency-bound simulation away from throughput parts. ECMP elephant flow analysis and per-port queue depth for fabric health.Which GPU should this job actually run on? Matched to what the job needs, not to whatever happens to be free.
Intelligence
Inference Runtime
vLLM · Triton · on-prem LLM
Designed to bind your in-perimeter model and surface which model runs on which GPU, at what speed — TTFT, ITL, KV-cache%, queue depth. Runtime targets include OpenVINO, AMD MIGraphX and Intel Habana. Zero cloud calls.Which model is running where, and how fast? Designed to bind to your own model, inside your own network. No cloud calls.
OversightGovernance Intelligence
RAG · IaC · audit-grade
Governance findings aligned to ISO 42001 and NIST AI RMF. Knowledge base grounded in your own documents, answered by the model you bind at deployment. Multilingual and per-PI signing are in development. IaC Connector showing live Terraform, Ansible, and Pulumi drift.Does our setup still match what we told the auditor? Findings in plain language — five languages in development — plus live drift in your infrastructure code.
Digital Twin
What-if · fault injection
Simulate cluster changes before committing them. Adjust GPU load, nodes, queue size, and memory pressure. Fault injection: node failure or degraded NIC. Recommended actions generated automatically.What happens if this node dies at 2am? Answer it on a Tuesday afternoon instead of during the incident call.
Infrastructure
Storage
Per-node I/O · NVMe health
Per-node read and write rates, filesystem utilisation, and NVMe health read from the drive’s own controller — not just throughput. The io-wait signal that separates an idle GPU from a storage-starved one comes off the same path.Is storage keeping up, and are the drives healthy? This is the difference between a GPU with nothing to do and a GPU waiting on a disk.
Nodes
Hardware registry · CMDB
Complete per-node hardware visibility: CPU, RAM, GPU (utilisation, VRAM, temp, power, PCIe TX, ECC, CUDA driver), storage, network, BMC. Living hardware registry — the CMDB your cluster never had.What hardware do we actually own, and what state is it in? The inventory your cluster never had, keeping itself current.
Topology
Topology canvas · RDMA
Topology canvas with rack-level drill-down. RDMA latency matrix: microsecond pairwise GPUDirect measurements (active probe, run on request). Five view modes: Cluster, Rack, Node, GPU Util, Network.How is all of this wired together, and where is it slow? A live map, down to the individual link.
Vulnerabilities
Offline CVE · signed bundles
Offline CVE scanning via signed .vulnbundle — designed for environments with no outbound connectivity. CVE × Pattern attribution. PDF export for compliance evidence.Are we exposed, and can we prove it without internet access? Signed offline scanning, with an export your auditor accepts.
Governance
Compliance Intelligence
Pramaan Score
Pramaan Score maps live infrastructure posture across the control catalogue — ISO 42001, NIST AI RMF, EU AI Act, DPDPA, CERT-In, SOC 2, GxP, MeitY, CAG, STQC. Gap engine with "Show me how" remediation.How close are we to the rules we are held to? One score from live readings off your own machines, and a ranked list of what to fix first.
Finance Intelligence
INR · kWh · kg CO₂
Your cluster cost INR 45,000 this month, across the jobs delivered — an illustrative figure from the lab reference cluster, not a customer result. Energy and carbon per job. PI-variant reports with student anonymisation. Pramaan evidence PDF for grant reporting.What did the cluster cost, and what did we get for it? Cost, energy and carbon per job — and how much of it was actually useful.
Admin & RBAC
13 roles · LDAP · PAM
13 user roles — admin, auditor, finance, pi-lead, cluster-engineer, researcher, supervisor, and more. LDAP-backed. All role changes are written to the hash-chain audit log.Who can see what, and who changed it? Thirteen roles tied to your existing directory, every change written to a tamper-evident log.
9
Telemetry layers
Job → Framework → Runtime → Container → Scheduler → Kernel → OS → Firmware → Hardware
255
Compliance probes built · 27 of 40 frameworks carry live evidence today
the rest say not assessed — honestly · on-premise · no outbound telemetry
13
User roles
admin · auditor · finance · pi-lead · researcher · supervisor · and more
Three solutions. One control plane.

Pick the problem you need to solve first.

Every Cognit deployment covers all three. These are the buyer entry points — where the pain is sharpest, the conversation starts.

Solution 01
Cluster Intelligence
See your GPU cluster as it actually is.
Pain
Grafana shows node-level utilisation. DCGM shows process metrics. Nobody can answer: why is this training run 30% slower than last week? Or: which team's job is burning GPU time right now? Investigations take days.
Need
GPU telemetry correlated from the Slurm job layer all the way down to kernel scheduling and NVLink state — in one dashboard, without building a fragile DIY stack or sending data to a cloud vendor.
What Cognit delivers
eBPF kernel instrumentation traces syscalls and scheduling events the kernel sees — not what a user-space exporter guesses. Eighteen tracepoints: file opens, device ioctls, reads, writes, io-wait, descriptor duplication, fork and exit. Anomalies the tracepoints witness are correlated to a job, a user, a node, and a timestamp. The investigation starts with the answer instead of a war room — you are reading evidence, not reconstructing it.
9
Telemetry
layers read
0
Data leaving
your network
0
Cloud
calls
Who it's for MLOps Engineer · Infra Lead · Platform Team · HPC Centre Manager · AI Programme Director
Solution 02
Cluster Efficiency
Attribute every GPU-hour. Recover the waste.
Pain
Finance asks for per-project GPU spend. The answer is a Slurm node-hour report nobody trusts. Teams over-provision jobs "just in case". Orphaned allocations sit idle for days. The CFO wants a number — not a cluster config dump.
Need
GPU-hours attributed to cost centres at cgroup granularity — using kernel-verified idle and holder detection joined to the device counters — not utilisation% alone. Finance-ready, not a raw Prometheus query that only the MLOps lead understands.
What Cognit delivers
Every GPU-second mapped to a Slurm job, user, cgroup, and cost centre — using kernel-verified idle detection. Chargeback reports export to CSV or your finance system. Right-sizing recommendations backed by 9 layers of evidence. We model the payback against your own utilisation data before you commit — and we do not quote a recovery percentage until we have seen your cluster.
~5.4%
Token-MFU (dense) at
100.0% reported utilisation
13
User roles
supported
0
Cloud
tax
Who it's for CFO / FinOps Lead · AI Programme Director · Research Head · Grant Manager · HPC Centre Manager
Solution 03
Cluster Governance
AI compliance — detected continuously, explained on-premise.
Pain
Compliance tooling typically means sending GPU telemetry to a third-party SaaS dashboard — violating the exact data residency obligations the CISO is trying to prove. Manual audit prep runs for weeks and still misses kernel-level drift events.
Need
Continuous governance monitoring — mapped to DPDP-India, ISO 42001, MeitY, and other frameworks — running entirely within your physical boundary. Findings explained in plain language, not raw log dumps, to be board-presentable.
What Cognit delivers
Every telemetry layer, every compliance check, and every LLM-generated finding explanation runs entirely within your perimeter. The Pramaan Score™ — a live composite governance metric across the control catalogue — is updated continuously. Audit evidence packages export in hours, not weeks.
40
Frameworks in
the catalogue
0
Verdicts moved by an
uploaded document
0
External
API calls
Who it's for CISO · Data Protection Officer · Compliance Lead · AI Programme Head · Board Risk Committee
Capabilities

Everything your cluster needs.
Nothing it doesn't.

Ten capabilities. One agent. Deployed in your environment — not ours.

See the Truth
Kernel-level attribution ties GPU allocations and activity back to the exact job and tenant. What DCGM, Slurm and Prometheus don't surface by default, Cognit correlates.
fd→GPU · per-tenant · io-wait
Know the Cost
GPU-hours wasted, energy and carbon (modelled), and cost — attributed per job, per user, per team. Your CFO finally gets a straight answer.
INR per job · kWh · kg CO₂
Stay Compliant
Pramaan Score maps your live infrastructure posture across ISO 42001, NIST AI RMF, EU AI Act, DPDPA, CERT-In, SOC 2, GxP and more — each gap ranked by how many controls it closes.
40 frameworks · one control catalogue
Runs in Your Jurisdiction
Fully on-premise. No outbound telemetry. Offline CVE bundles with hash-chain integrity verification. No foreign-jurisdiction cloud dependency — the deployment sits inside your own boundary, on your own hardware.
On-premise · no egress · your boundary
Detect Failures Early
Pre-built detection patterns across scheduling, memory, thermal, hardware, storage, network, jobs, and accounting. Ships ready — no configuration required.
Pattern library · auto-detect
Govern Workloads
Per-job Pramaan Trust Score is designed to flag workloads at admission. In development — declaration capture and drift detection ship today. Declarations drive policy; the kernel agent drives proof.
Declaration capture · drift detection · hash-chain audit
Pattern Detection Library
From GPU Idle Despite Queue to InfiniBand Link Degradation — each pattern is independently tracked, status-flagged (ACTIVE / WARNING / CLEAN / PENDING), and linked to causal explanation.
Status-tracked · scheduling to network
GPU Routing Intelligence
Recommends the right GPU for each workload by job class — LLM inference to HBM2e, CUDA training to CUDA-optimised silicon. Validated on NVIDIA today; the kernel evidence layer is vendor-agnostic by design, with AMD next.
Workload routing · validated on NVIDIA · AMD next
Compliance Intelligence
Not a checklist — a live gap engine. Each gap shows which frameworks it touches, which controls it closes, the exact remediation step, and API endpoints to verify after the fix.
Gap engine · cross-framework · Show me how
Built for Your Requirements
Every cluster is different. If your environment needs a capability Cognit doesn't ship today — a custom integration, a specific compliance framework, a hardware connector, or a bespoke reporting module — we build it with you.
Custom integrations · bespoke modules · your stack · your rules
Custom scheduler connector
Proprietary hardware telemetry
Bespoke compliance framework
Custom reporting & dashboards
Internal audit format export
Your use case →
Tell us what you need →
What your current stack misses

Four ways dashboards lie. Every day.

Your GPU monitoring stack gives you device-level aggregates. It cannot tie kernel-level evidence back to the job or the tenant. That gap is the whole problem.

util% = useful work
→ MFU tells the truth
A GPU pinned at 95% can be spinning, waiting on an NCCL collective, or KV-cache thrashing. DCGM sees the chip is busy. It does not see whether the work is real. MFU — model FLOP utilisation — does.
We measured this on our own run
100.0% utilisation · ~5.4% token-MFU
L40S, Qwen2.5-14B-AWQ on vLLM, 32 concurrent streams, 20,288 generation tokens over 61s. A decode-heavy run is memory-bandwidth-bound, so single-digit MFU here is the physics ceiling, not proof of waste — the waste question is asked of prefill and training phases. The point stands either way: the gauge read 100% and could not tell you which.
Token-throughput estimate, not native MFU — 2·N·tokens/s ÷ peak FLOPS, against the dense BF16 peak. Against the 362 TFLOPS sparsity peak the same run reads 2.7% — we quote dense because vLLM cannot reach the sparsity figure, so that denominator flatters nobody honestly. AWQ is 4-bit, so this compares quantised throughput to a dense peak. Synthetic load, not a customer workload. 15 Jun 2026; raw scrapes retained.
idle% = available
→ io-wait tells the truth
A GPU at 20% during AlphaFold's MSA phase isn't free — it's blocked on Lustre reads. Reclaiming it would kill the job. Kernel io-wait is what tells idle apart from storage-starved.
GPU watts = what it costs you
→ the rail tells the truth
Driver-reported power is the accelerator alone. Your bill, your PUE return and your energy report are all the node — CPU, memory, fans, drives and supply losses included. Read one and infer the other and you will be wrong by most of the number.
We measured this on our own cluster
15.8 W at the driver · 142 W at the node rail
The same idle GPU, read the same minute. The driver is not lying — it is answering a narrower question than the one the invoice asks.
NVML per-device power against the node’s own DCMI rail via /sys/class/hwmon. These measure different things by design; that is the point. Single node, idle. 10 Aug 2026.
shared = attributed
→ kernel fd-table tells the truth
On a shared, MIG, or opaque GPU, who is actually holding the allocation? The kernel's file-descriptor table names the tenant — and it is the record that survives, because no downstream Prometheus join can reconstruct the name once the aggregate is emitted.
Use Cases

What the kernel layer
changes in practice

Three real workloads where the standard stack gives you a number. Cognit gives you the reason.

USE CASE 01
vLLM inference — TTFT spiking
DCGM: 90% util. Problem unseen.
KV-cache is full. vLLM is preempting and swapping. The GPU looks busy because it is — doing the wrong work. Kernel attribution names the tenant causing the evictions. Add capacity or evict the right job. Don't trust the util gauge.
USE CASE 02
AlphaFold — the GPU that looks idle
DCGM: 20% util. Operator reclaims it.
MSA search is running: massive sequential reads from BFD and UniRef on Lustre. The GPU is waiting on storage. Kernel io-wait proves it's blocked, not free. The fix is the data path — not handing the GPU to another job.
USE CASE 03
Isaac Sim — the H100 that lost to a gaming card
Scheduler: GPU assigned. Job runs. Meter reads busy.
We measured it. On physics-only training an H100 ran ~3,800 steps/s against ~8,700 on a consumer RTX 4060 — the task is latency-bound, so the H100's throughput is wasted while the utilisation meter shows it busy the whole time. Turn a camera on and the H100 does not degrade quietly: it hard-crashes with a named signature, VK_ERROR_DEVICE_LOST — the renderer requires RT cores, and the H100 has none. Bigger GPU ≠ faster. Measured 21 Jun 2026 on real hardware. The agent already reads the card model and derives its RT-core capability, so it detects the mismatch — H100, A100 and V100 have no RT cores; L40S, A10 and RTX do. Gating placement on that verdict is in development — today the placement gate reads the declared job class.
Field Results

What we've already caught.

Short versions only — each of these is measured, dated, and from a real machine: a customer's fleet or our own reference cluster. Ask for the long version on a call.

CAUGHT 01
Training detected the moment it started
Customer fleet · 11×RTX 4090 · July 2026
The fleet heatmap picked up a training run starting on GPUs 4–7 with zero configuration — no agent tuning, no job hooks — while live inference findings kept firing on GPUs 0–3. The operator didn't tell us; the kernel did.
CAUGHT 02
7 of 11 GPUs stranded on first look
Same fleet · first snapshots · June 2026
Seven GPUs fully idle, and the other four holding loaded models at near-zero utilisation, plus a model-placement warning — on a box the dashboards called healthy. 64% of the fleet — seven full GPUs — was reclaimable capacity.
CAUGHT 03
The node that flagged itself at 79°C
While recording a customer demo · July 2026
A thermal warning fired at 79°C on a real node mid-recording — unscripted, unstaged, surfaced before anyone asked. 79°C is not exotic; the point is that nobody staged the moment. The best product demo we never planned.
CAUGHT 04
8,356 MiB of VRAM held at 0% utilisation
Our own reference cluster · August 2026
The stranded-capacity detector caught a process holding more than 8 GiB of VRAM while doing nothing with it — waste a utilisation gauge cannot see, because idle VRAM reads as a quiet, healthy GPU. Named the holder from the kernel's own fd table.
Workload Identity

It says python3.
We say GROMACS.

Renaming a process changes what it calls itself. It does not change what the kernel executed. Cognit identifies named scientific and AI workloads from kernel evidence — and is building the attestation layer that makes that identity something a tenant cannot forge.

14 — named signatures built
Workload fingerprints spanning molecular dynamics, DFT, CFD, structural FEA, genomics, cryo-EM, neuroimaging, quantum simulation, inference serving, seismic imaging and weather models. Verified signatures only — the library ships nothing speculative.
6 — verified firing on live runs
GROMACS, Quantum ESPRESSO, Ansys Fluent, Abaqus, CP2K and Triton have each fired on live runs on our own cluster — real binaries on real kernels, not fixtures.
4 — kernel-attestable today
For GROMACS, Quantum ESPRESSO, CP2K and Triton, the kernel's own record of the executed binary can contradict a renamed process: exec -a changes what a process calls itself; it cannot change the inode the kernel ran. Full identity attestation across the library is in development.
The signature library — 13 tools, 14 fingerprints (Quantum ESPRESSO carries two) — molecular dynamics to weather models
GROMACS
Quantum ESPRESSO
CP2K
VASP
Ansys Fluent
Abaqus
NVIDIA Triton
NVIDIA Parabricks
RELION
FastSurfer
CUDA-Q
Seismic RTM / FWI
Weather / NWP

Wrapper-launched tools — Fluent, Abaqus, Parabricks — are recorded as exactly that: a wrapper legitimately shows an interpreter to the kernel, so we say identified, not attestable rather than pretending. The evidence ladder ranks every identity source by how hard it is to forge, and says so on the record.

How findings are made

Deterministic by design.
No model decides a verdict.

Every finding is produced by symbolic rules over kernel evidence — the same inputs produce the same verdict, and verdicts carry the witnesses that produced them. A language model writes the plain-English explanation. It never decides the finding.

Symbolic, not statistical
Pattern verdicts come from a rules engine with named conditions and thresholds — not from a black-box model whose reasoning cannot be replayed. An auditor can read the rule, read the evidence, and re-derive the verdict by hand.
Cited, chained, replayable
Findings carry their evidence: the counters read, the thresholds crossed, the kernel records that witnessed them — appended to a tamper-evident hash chain. A verdict without evidence attached is a bug, and we treat it as one.
Guardrails that can say no
The UI cannot render "measured" without an explicit synthetic:false from a real source. The reference board hides by default anything unverified. An uploaded document can never move a verdict. The system is built to refuse a claim it cannot stand behind. Each of these is enforced in code and pinned by tests — a violation is a defect, and we say so.
Deployments can run with the language model switched off entirely — findings unchanged, only the prose gets shorter. AI narrates. It never adjudicates.
Pramaan Score

Governance that's
measured, not claimed.

Forty frameworks in the control catalogue, assessed from live infrastructure evidence — not self-assessment checkboxes. Coverage depth is reported per control in the work paper, from the evidence present in your estate at the time of the run. Every gap is ranked by cross-framework impact with a clear remediation path.

NC
Governance policy document missing. Closes 10 controls across CERT-In CI-5.1, DPDPA 1.2/1.5, EU AI Act AIA-2.1, ISO 42001, NIST AI RMF.
OFI
No CVE scan recorded. Ingest a signed .vulnbundle and run a scan. Closes CERT-In CI-2.1, NIST MAN-3.1.
C
System operations monitoring — Conformant. Hash-chain audit log + Prometheus telemetry + vulnerability intake all verified.
40
Frameworks in the control catalogue
A panel reads “measured” only when the payload carries an explicit synthetic:false from a real source — a failed or empty read renders “unavailable”, never a green badge. And the board hides by default. A card is shown only if it names a real node and carries a real measurement, or is honestly empty — everything else is hidden rather than dressed up. On the default view of our own reference board, with that mode on, 15 of 39 cards render and 24 are hidden, including our own per-job trust score. Measured 10 Aug 2026. Panels currently modelled rather than measured: energy and inference. A qualified auditor must review before reliance.
FRAMEWORKS COVERED
ISO 42001
NIST AI RMF
SOC 2 TSC (mapping)
EU AI Act
DPDPA 2023
CERT-In 2025
GxP / 21 CFR 11
MeitY 2025
CAG IS Audit
STQC AI QA
40 frameworks · every gap ranked by cross-framework impact
How we engage

Start with a fixed-scope study.
Everything credits forward.

We don't ask for a platform commitment before you've seen your own numbers. Each step is fixed scope, fixed price, and credited in full against the next — so the only thing you risk is the first one.

+ 18% GST applies on INR invoices
Step 01 · 1–2 weeks
Scoping Study
You send metrics. We tell you where the waste and the risk are — before anything touches your environment.
INR 2,50,000 fixed
Credited in full against Step 02
  • Export from Prometheus, DCGM or Slurm accounting — a screenshot is enough to start
  • Written findings: where GPU-hours are going and which patterns are firing
  • Indicative recoverable headroom for your cluster
  • No agent, no access, nothing installed
Start here
Step 03 · scoped per engagement
Full Audit
The complete governance engagement across your estate, ending in a certification-body-style work paper.
from INR 15,00,000
Scoped to cluster size and framework set · Proof of Audit credited
  • Full control catalogue — 40 frameworks, with per-control provenance showing which carry live evidence in your estate
  • Ranked remediation with cross-framework impact and the API endpoint to re-verify
  • Auditor work paper with per-control provenance — live evidence vs attached document
  • Findings the platform did not measure are marked not assessed, never scored
Talk to us
AFTER THE AUDIT — ONGOING PLATFORM

Continuous observability and the compliance plane, deployed on your hardware — Observe for attribution and efficiency, Govern for the full compliance plane. Licensed per GPU, annually, and scoped to your cluster. Pricing on request — we quote it against what the audit found, not from a table.

Volume tiers apply above 64 GPUs and aggregate across clusters. Sovereign and defence deployments are scoped separately — talk to us.

Prices are set independently in each currency and hold for the term of the engagement — they are not converted at a daily rate. Revenue figures reported elsewhere on this page are stated in the currency actually received. Formal quotations are currently issued in INR; a quotation in your selected currency is available on request.

Architecture

Cognit reads every layer.
Most stacks skip the kernel band. That's where we start.

Most stacks bolt one tool per band and call it coverage. The runtime and device ends are well-observed. The kernel — where causation actually lives — is a blind spot.

Layer
What it tells you
Evidence path
Dependency
Job / Workload
State, runtime, GPU alloc, exit code
your scheduler's accounting
◆ integrated
Framework
Loss curves, tokens/s, samples/s
your training profiler
◆ integrated
Runtime / Serving
TTFT · ITL · KV-cache% · MFU
your serving runtime's metrics · partial
◆ integrated
Container
Image provenance, cgroup pressure, per-cgroup attribution
/sys/fs/cgroup — kernel
● kernel-direct
Scheduler
Queue wait, placement, GRES allocation
your scheduler's queue
◆ integrated
Kernel (Linux)
io-wait · fd→GPU map · per-tenant attribution · declared-vs-detected behaviour
eBPF — Cognit's agent
★ ours
OS Userspace
Device counters, NCCL events, dmesg, clock sync + offset
chronyc · timedatectl · vendor SMI
● kernel-direct
Firmware / BMC
Power, thermal, fan, PSU, ECC, per-device watts
/sys/class/hwmon · Redfish/IPMI
● kernel-direct
Hardware
GPU util%, temp, NVLink, IB counters, PCIe link gen/width
/sys/bus/pci · vendor SMI · fabric counters
● kernel-direct

★ ours — instrumentation we wrote and ship. ● kernel-direct — at least one signal at this layer is read from the kernel's own interface with no vendor daemon in the path; the evidence column names which, and what still comes through a vendor tool. ◆ integrated — read from the system of record, because Slurm is the authority on job state and asserting otherwise would make the evidence worse, not better.
Tools named are examples, not requirements. The four kernel-direct rows read identically whatever scheduler or GPU vendor you run — that is the point of reading the kernel.Nine layers read · one ours · four with a kernel-direct path · four from the source of record.

Who We Serve

Built for people who run real clusters.
Not cloud-credit holders.

If you own the hardware, share GPUs across teams, and have ever wondered where the hours actually went — this is for you. Five situations, not five industries — and in every one of them, a gauge is lying to you.

Neoclouds & GPU providers
Utilisation lies
You bill by the GPU-hour, and you cannot prove what the tenant actually got. The kernel's file-descriptor table names the tenant holding a shared allocation and survives the aggregation that erases it everywhere else — attribution you can finally take to billing, instead of writing the hours off.
→ You are the ideal customer if you have ever refunded a tenant for a GPU you could not account for
Life sciences & pharma R&D
Starvation lies
AlphaFold and molecular-dynamics pipelines are I/O-starved by design. Your GPUs read 20% and your infrastructure team wants them back — but reclaiming a storage-blocked GPU kills the run. GxP and 21 CFR Part 11 evidence comes off the same trail.
→ You are the ideal customer if someone has proposed reclaiming a "quiet" GPU mid-pipeline
Enterprise AI & ML platform teams
Chargeback lies
One shared cluster, many teams, and a chargeback model nobody trusts. The question that never gets a clean answer is "whose job is this, and did it earn its hours?" — because the aggregate is emitted before the tenant is ever named.
→ You are the ideal customer if your GPU chargeback report is argued with every quarter
Financial services
Cost reports lie
Heavy inference and risk workloads under strict audit. Cognit attributes GPU cost to teams and jobs at kernel granularity, and the compliance evidence — mapped to DPDPA, ISO 27001 and SOC 2 (TSC mapping) — comes off the same trail as the efficiency numbers. One instrument, both answers.
→ You are the ideal customer if model-risk or internal audit has asked who ran what on the GPU estate — and the answer took days
Robotics & embodied AI
Uptime lies
Simulation workloads are latency-bound, not throughput-bound, and the biggest accelerator in the rack is often the wrong one. We measured an H100 losing to a consumer RTX 4060 on the same physics task — and a utilisation meter showed that H100 busy the whole time.
→ You are the ideal customer if your sim fleet mixes datacentre and workstation GPUs
Not for you if —
You call a hosted LLM API · you run a single-user box · you use managed serverless GPUs · your infrastructure is 100% hyperscaler-managed. No shame — just not the problem we solve.
Product-Market Learning — what early conversations taught us

Detection is a commodity.
Evidence is not.

Four objections from early conversations — paraphrased, not verbatim — and what each one taught us. We publish these because the people who say no teach you what the product is.

TOOLING BUILDER
A sysadmin — or an AI agent — can find an idle GPU.
Correct. Detection alone is worth $0. We stopped selling detection. What survives is what detection cannot produce: signed, chained, audit-grade evidence of who ran what.
GOVERNMENT BUYER
Why pay for what agents can do a basic version of?
Because a basic version cannot sign, chain, or survive an audit. Price follows proof, not effort.
EDUCATION INSTITUTE
No one holds us accountable for efficiency.
No accountability → no budget. Accountability is now our first qualifying question.
RELIABILITY ENGINEER
This is preventive maintenance — a known space.
If we lead with health, we compete with maintenance. We lead with attribution + evidence.
The pattern in every no: people who don't have to prove anything won't pay. The accountable party will. Our ideal customer: operators selling compute with SLAs, and AI teams facing a regulator or an auditor.
Design Partners — everyone who said yes must prove something

Five design partners.
Five GPU verticals.

None of them is named here — partners get named when they say so, not before. Each engagement exists to prove one specific thing.

live deployment  ·  in progress

EU inference provider
11×RTX 4090 · vLLM in production
Licensed deploy since June 2026. The fleet heatmap picked up training starting on GPUs 4–7 with zero configuration; live inference findings on GPUs 0–3. He has corrected our detection rules from his own physics — twice. We publish the corrections.
Robotics / Isaac Lab engineer
sim-to-real RL · Apptainer on HPC
Two paid milestones delivered. His field report handed us seven real failure modes — with crash signatures — that standard monitoring stacks miss today. That is our robotics detection catalogue.
Crypto / prediction-market quant
CUDA optimisation · latency-bound
Currently auditing Cognit against his own stack — an adversarial user by profession, which is exactly the reviewer we want.
GPU optimisation services provider
benchmarking & profiling practice
Testbed partner for accuracy validation: their profiling ground-truth against our attribution, on real client-style workloads.
Edge AI / ARM engineer
MediaPipe GPU on ARM boards
Scoping the edge story: the same evidence layer on RK35xx-class hardware.

Also in motion: an ongoing conversation with a major Indian GPU cloud · an early conversation with a global pharma manufacturer (GxP).

About

Built in India.
Runs in your jurisdiction.

Cognit is an India-based infrastructure software company building the governance and observability layer for on-premise AI compute. We exist because the organisations that need this most — GPU clouds, enterprise AI platform teams, research and life-sciences computing — cannot send cluster telemetry to a foreign-jurisdiction SaaS dashboard, and no one was building for them.

We fund the deep-tech build with a working services practice — embedded systems, cloud communications and cybersecurity — which delivered and collected ₹13.3 lakh in FY 2024-25. Self-funded, debt-light and audited. We move carefully, we say what we mean, and we build in public where we can.

Bangalore, India
Self-funded · revenue-backed
On-premise · No egress
Independent · India-registered
REGISTERED & CERTIFIED
DPIIT DIPP169103
ISO/IEC 27001:2022 QCC/E808/1025
Udyam MSME KL-10-0040651
Kerala Startup Mission Innovation Grant, approved Jul 2026
2 patents filed
Two patents filed — cluster management via telemetry correlation, and boot-time firmware-tamper attestation.
9
Telemetry layers read — Job to Hardware
117
BMC sensors read on the lab reference node — timestamped, per sensor
40
Frameworks in the control catalogue — coverage reported per control, not summarised
255
Compliance probes built — 27 of 40 frameworks carry live evidence today; the rest say not assessed
Get in touch
sales@cognit.run
For pilots, partnerships, and early access conversations.
The Team

The people building this.

A working team, not a founder and a deck. Engineers who run the lab cluster, design the compute module and deliver on customer sites — and specialists we bring in for the parts that need a specialist: Rust systems and performance, hardware verification, audit, and robotics.

Krishnadas Puthukudy
FOUNDER
Capital, operations and long-term strategy. Heads the services business that funds the build — selling to and delivering for multinational industrial customers, including two turnkey PCB programmes for a global connector and sensor manufacturer, both delivered and signed off by the customer. Previously Saudi Chevron.
Agnidipa Manna
CARRIER BOARD ENGINEER
Designs the COM-HPC carrier board — the indigenous compute module.
Sainath Reddy
LAB CLUSTER ENGINEER
Runs the lab cluster the product is developed against. FPGA and VLSI background.
Savin Sundar
FORWARD DEPLOYMENT ENGINEER
PMP, PMI-ACP and CSM certified. Led multi-country ERP programmes across India and Africa to on-site go-live under hybrid Agile–Waterfall governance. SAP MM consulting before that, and project controls in oil and gas before that again.
Kajal J.
GROWTH & PRODUCT MARKETING
Six years in B2B sales and go-to-market for early-stage startups across APAC and EMEA. Has run outbound and qualification for an AI-observability platform — the same buyers this one has: AI engineering teams, MLOps leads and compliance owners. Owns positioning, segmentation and messaging.
Yoshita Pacholy
OPERATIONS
Project delivery and programme management. Previously Xieno, Ericsson.
Parag Somani
STRATEGY
Technology investment banker and former founder.
Sandeep Krishna
CO-FOUNDER
Network engineering. Previously Telstra and Ericsson.
ADVISORS & SPECIALISTS
Bobby Mathews
RUST SYSTEMS & PERFORMANCE
Systems and performance engineering in Rust — executor starvation, tail latency, backpressure and lock-free real-time paths. Publishes original benchmark research, including a measured 101× p99 scheduling-overhead increase when blocking work saturates an async runtime. GPU-accelerated processing of high-frequency sensor data, and compute and shader tooling over Vulkan and WebGPU.
Saheb Mandal
HARDWARE
COM-HPC carrier design verification — power stage, eFuse telemetry and PMBus.
Shruthi P Nair
AUDIT & COMPLIANCE
Audit, compliance and cloud governance. Reviews the work-paper format against how an auditor actually reads it.
Ebin Babu
LINUX SYSTEMS
Linux systems engineering.
Paul de Sainte Agathe
vLLM & ML PERFORMANCE
Runs the multi-endpoint vLLM deployment the runtime telemetry is validated against.
Truong N.
ROBOTICS · ROS2
Isaac Sim and Isaac Lab on scheduler-managed clusters. Ran the RT-core experiment.
Get Started

Ready to see your cluster
as it actually is?

Share a few details and we'll get back to you within 2 business days. No demo-bot. No automated funnel. A real conversation with the team.

30-minute technical walkthrough on our lab node
We review your metrics and tell you where the waste is — before touching anything
Read-only pilot on your cluster · nothing leaves your network
Findings from your own exported metrics — nothing installed to start
No commitment required to get started

By submitting, you agree we'll use this info to get in touch. That's it.
or
sales@cognit.run

Message sent.

We'll be in touch within 2 business days at the email you provided.