Explore your site
DAIT / Monitoring & operations
Monitoring & operations

We watch the systems — and we call the people

Deployed compute only earns if it stays up. Node monitoring and incident response are delivered through DAIT Global Solutions and its IT operations partner — a 24/7 practice spanning infrastructure, databases, services and network, with one detection loop, named owners, and a human on the phone the moment something breaks.

24/7Monitoring & on-call
4 hoursImmediate-priority SLA
Under 15 minTriage window
The DAIT Customer Portal

Everything about your node, in one place

Hosts and compute customers sign in to a single portal. It shows the state of every node on your site and what it has produced — no request to make, no report to wait for.

Node status

Live health per unit

Each deployed node on your site, its current condition and whether anything is being worked on. If an incident is open, you can see it rather than hear about it later.

Power and battery

Electricity and storage

Electricity drawn and the state of the battery stack, so the resilience the node adds to your building is visible rather than assumed.

Earnings

Revenue from your site

What the node has generated, tied to the site it sits on. Hosts see what their space is earning without asking anyone.

The portal is available to site hosts and compute customers once a node is live. DAIT is at pre-pilot stage, so no node is in commercial operation yet.

What we monitor

Full-stack coverage on every host

Four layers are instrumented continuously on each machine — from kernel counters to whether a public endpoint actually answers.

Host level

Infrastructure

CPU, memory, swap, disk capacity and I/O, and uptime — core health at a glance, per host.

Data layer

Databases

PostgreSQL, MSSQL and Oracle — connection counts, query performance and availability, watched as a first-class layer rather than inferred from host load.

Application layer

Services

Application and system services on each host, including containers: state, health check, restart count and start type.

Network

Interfaces & links

Throughput, traffic, queue length and physical link status on every NIC — management, storage and guest interfaces alike.

Depth of instrumentation

Every panel breaks down to mean, last, max and min

A single number rarely explains an incident. Each resource is read through four statistical lenses and sliced by the thing that actually caused it — process type, mount point, or interface.

LensWhat it tells us
MeanRolling average over the trend window. Exposes slow drift — memory creep, gradual CPU saturation.
LastThe live value that triggered the view. What the operator sees first.
MaxHighest spike in range. Correlates against batch jobs, backups and burst traffic.
MinLow watermark. Confirms recovery after a fix and shows the real headroom.
CPU

By process type

System, user, iowait and steal are separated, so time spent is attributed rather than guessed at.

Storage

Performance and capacity

Read/write throughput, utilization and queue depth per device; filesystem space and free inodes tracked separately, with early headroom warnings.

Network

Per device, not aggregate

Metrics are held per interface — physical NICs, loopback, bridges and tunnels — alongside carrier state, LACP status and flap detection.

Beyond the host

A host being up is not enough — the service has to answer

External checks

Endpoint reachability

Public endpoints are probed on schedule from outside the network — the same vantage point as a real user — for HTTP status code and body match, not just ping.

  • Status code and content match on a fixed interval
  • SSL expiry, certificate chain and revocation tracked
  • Response time held as p50 and p95, with error rate
Virtual layer

Hypervisor and guests

Physical host and virtual layer are watched in one view rather than in separate tools, so triage does not start with tool-hopping.

  • Hypervisor: management daemon health, CPU ready, ballooning and HA state
  • Datastore: capacity, free inodes and I/O latency per LUN
  • Guests: per-VM resource use tied back to the host
Access, identity and audit

Named logins. Scoped roles. Full attribution.

Every team member signs in under their own account before touching monitored data. There are no shared credentials, so every action is tied to a person, a role and a timestamp.

One person, one account

Individual identity

A personal account is required for the DAIT Customer Portal and for internal systems. Multi-factor authentication is enforced and the session is logged.

Least privilege

Role-based access

L0, L1, DBA and DevOps roles each see only what the role needs, with isolation applied at the data layer. Access is reviewed and revoked automatically on role change.

Audit trail

Every action logged

Acknowledgements, comments and configuration changes are audit-trailed. The full history is exportable and can be shared on request.

Alert detection and escalation

From trigger to closure, in five steps

Every threshold breach follows the same path. No guesswork about who owns it, and no waiting for office hours to reach a specialist.

1

Alert triggered

Threshold breached. Dashboards flash and email fires instantly.

2

Initial triage

Classified as a routine L0/L1 case or a complex database case. Decided inside the 15-minute triage window.

3

Resolve or escalate

L0 and L1 issues are fixed directly. Anything more complex pages the on-call DBA or DevOps engineer.

4

Customer informed

A phone call goes out as soon as the alert is received — at the start, not after the fix.

5

Resolution and closure

Final update shared, timeline archived, history sealed and auditable.

Ticket lifecycle and SLA

Auditable from open to close

Timestamps, confirmations and reopens are automatic, so there is no manual step to forget. Immediate-priority tickets carry a four-hour SLA that starts the moment the ticket opens.

1

Opened

Source alert is linked to the ticket. On Immediate priority the SLA clock starts instantly.

2

Assigned

A named owner is set with a timestamp. The queue is visible to the operations team and to the customer.

3

Resolved

Fix applied and an automatic confirmation email is sent.

4

Reopen on reply

Any customer reply reopens the ticket automatically. Nothing gets closed by silence.

5

Closed and surveyed

A satisfaction survey goes out 24 hours after closure. The history stays searchable.

Priority, SLA status and effort are tracked on every row of the open queue, in the same view for the operations team and the customer.

Communication

A human voice at every step

Alerts land on dashboards, but customers hear a person — not just a ticket number.

01

Immediate first call

The customer is phoned as soon as the alert is received, before triage has finished.

02

Progress updates

Status is shared while the issue is being worked. No silent queues.

03

The right specialist, any hour

Escalation reaches the on-call DBA or DevOps engineer directly, without waiting for office hours.

04

Confirmed closure

A final call confirms the fix, the report is attached, and the survey follows 24 hours later.

Ask how a node would be monitored at your site

Tell us what you run and where. We will walk you through the coverage, the alerting thresholds and the escalation path that would apply to your deployment.

Talk to operations