We watch the systems — and we call the people
Deployed compute only earns if it stays up. Node monitoring and incident response are delivered through DAIT Global Solutions and its IT operations partner — a 24/7 practice spanning infrastructure, databases, services and network, with one detection loop, named owners, and a human on the phone the moment something breaks.
Everything about your node, in one place
Hosts and compute customers sign in to a single portal. It shows the state of every node on your site and what it has produced — no request to make, no report to wait for.
Live health per unit
Each deployed node on your site, its current condition and whether anything is being worked on. If an incident is open, you can see it rather than hear about it later.
Electricity and storage
Electricity drawn and the state of the battery stack, so the resilience the node adds to your building is visible rather than assumed.
Revenue from your site
What the node has generated, tied to the site it sits on. Hosts see what their space is earning without asking anyone.
The portal is available to site hosts and compute customers once a node is live. DAIT is at pre-pilot stage, so no node is in commercial operation yet.
Full-stack coverage on every host
Four layers are instrumented continuously on each machine — from kernel counters to whether a public endpoint actually answers.
Infrastructure
CPU, memory, swap, disk capacity and I/O, and uptime — core health at a glance, per host.
Databases
PostgreSQL, MSSQL and Oracle — connection counts, query performance and availability, watched as a first-class layer rather than inferred from host load.
Services
Application and system services on each host, including containers: state, health check, restart count and start type.
Interfaces & links
Throughput, traffic, queue length and physical link status on every NIC — management, storage and guest interfaces alike.
Every panel breaks down to mean, last, max and min
A single number rarely explains an incident. Each resource is read through four statistical lenses and sliced by the thing that actually caused it — process type, mount point, or interface.
| Lens | What it tells us |
|---|---|
| Mean | Rolling average over the trend window. Exposes slow drift — memory creep, gradual CPU saturation. |
| Last | The live value that triggered the view. What the operator sees first. |
| Max | Highest spike in range. Correlates against batch jobs, backups and burst traffic. |
| Min | Low watermark. Confirms recovery after a fix and shows the real headroom. |
By process type
System, user, iowait and steal are separated, so time spent is attributed rather than guessed at.
Performance and capacity
Read/write throughput, utilization and queue depth per device; filesystem space and free inodes tracked separately, with early headroom warnings.
Per device, not aggregate
Metrics are held per interface — physical NICs, loopback, bridges and tunnels — alongside carrier state, LACP status and flap detection.
A host being up is not enough — the service has to answer
Endpoint reachability
Public endpoints are probed on schedule from outside the network — the same vantage point as a real user — for HTTP status code and body match, not just ping.
- Status code and content match on a fixed interval
- SSL expiry, certificate chain and revocation tracked
- Response time held as p50 and p95, with error rate
Hypervisor and guests
Physical host and virtual layer are watched in one view rather than in separate tools, so triage does not start with tool-hopping.
- Hypervisor: management daemon health, CPU ready, ballooning and HA state
- Datastore: capacity, free inodes and I/O latency per LUN
- Guests: per-VM resource use tied back to the host
Named logins. Scoped roles. Full attribution.
Every team member signs in under their own account before touching monitored data. There are no shared credentials, so every action is tied to a person, a role and a timestamp.
Individual identity
A personal account is required for the DAIT Customer Portal and for internal systems. Multi-factor authentication is enforced and the session is logged.
Role-based access
L0, L1, DBA and DevOps roles each see only what the role needs, with isolation applied at the data layer. Access is reviewed and revoked automatically on role change.
Every action logged
Acknowledgements, comments and configuration changes are audit-trailed. The full history is exportable and can be shared on request.
From trigger to closure, in five steps
Every threshold breach follows the same path. No guesswork about who owns it, and no waiting for office hours to reach a specialist.
Alert triggered
Threshold breached. Dashboards flash and email fires instantly.
Initial triage
Classified as a routine L0/L1 case or a complex database case. Decided inside the 15-minute triage window.
Resolve or escalate
L0 and L1 issues are fixed directly. Anything more complex pages the on-call DBA or DevOps engineer.
Customer informed
A phone call goes out as soon as the alert is received — at the start, not after the fix.
Resolution and closure
Final update shared, timeline archived, history sealed and auditable.
Auditable from open to close
Timestamps, confirmations and reopens are automatic, so there is no manual step to forget. Immediate-priority tickets carry a four-hour SLA that starts the moment the ticket opens.
Opened
Source alert is linked to the ticket. On Immediate priority the SLA clock starts instantly.
Assigned
A named owner is set with a timestamp. The queue is visible to the operations team and to the customer.
Resolved
Fix applied and an automatic confirmation email is sent.
Reopen on reply
Any customer reply reopens the ticket automatically. Nothing gets closed by silence.
Closed and surveyed
A satisfaction survey goes out 24 hours after closure. The history stays searchable.
Priority, SLA status and effort are tracked on every row of the open queue, in the same view for the operations team and the customer.
A human voice at every step
Alerts land on dashboards, but customers hear a person — not just a ticket number.
Immediate first call
The customer is phoned as soon as the alert is received, before triage has finished.
Progress updates
Status is shared while the issue is being worked. No silent queues.
The right specialist, any hour
Escalation reaches the on-call DBA or DevOps engineer directly, without waiting for office hours.
Confirmed closure
A final call confirms the fix, the report is attached, and the survey follows 24 hours later.
Ask how a node would be monitored at your site
Tell us what you run and where. We will walk you through the coverage, the alerting thresholds and the escalation path that would apply to your deployment.
Talk to operations