Project 05

HydroSentinel AI

Project built during the USAII Global AI Hackathon 2026

Starting question

Can a hidden water leak inside a building be detected before visible damage appears, without manually inspecting every pipe?

Context-aware water anomaly and leak-detection decision support using synthetic telemetry.

View Honor

Public demo interface
HydroSentinel public demo showing scenario selection and analysis

Why we built HydroSentinel

WHY?

HydroSentinel started during a seven day AI hackathon. My team and I joined because we wanted to find a real problem with a negative impact and try to build something useful around it. We spent roughly five hours researching and discussing problems that mattered, while also asking what we could realistically address with an AI prototype in seven days. We kept coming back to water loss in school facilities.

What interested us was the detection problem. A maintenance manager can repair a leak once it is known, but cannot inspect a large network of hidden pipes every day. We wanted to see whether flow and pressure data could help identify unusual behavior early enough to justify an inspection.

WHY?

We did not have access to live telemetry from an instrumented school, and collecting real sensor history and verified leak events was not possible during the hackathon. We created synthetic CSV telemetry instead, including normal operation, high activity periods, pressure variation, noise and injected leak events. The first prototype used Python, Streamlit and Isolation Forest.

Testing showed that high water use did not always mean a leak. Events and high occupancy could create legitimate increases in demand, so we added operating context to the analysis.

After the hackathon, we kept developing HydroSentinel. The current version is a web application with a diagnostic classifier, a separate loss estimator, learned no-leak baselines, a FastAPI backend and a Next.js frontend. It also includes validation, testing, explainability and a public demo.

Technology

Data

The backend accepts four telemetry fields:

  • Timestamp
  • Flow_Rate_LPM
  • Avg_Pressure_PSI
  • Occupancy_Status
WHY?

It checks that the required fields are present. Flow must be numeric and between 0 and 500 L/min. Pressure must be numeric, greater than 0 and no more than 150 PSI. Occupancy_Status must be one of Class_Hours, After_Hours, Vacation or Event. Invalid rows are removed, and if no usable rows remain, analysis stops.

The current diagnostic features are:

  • Flow_Rate_LPM
  • Avg_Pressure_PSI
  • Occupancy_Status
  • Hour
  • Is_Weekend

Hour and Is_Weekend are derived from Timestamp. Numeric features pass through directly, and Occupancy_Status is encoded with OneHotEncoder(handle_unknown="ignore").

Model

There are two learned tasks. Classification and loss estimation are separate models.

Classification

RandomForestClassifier

Settings
n_estimators = 250, random_state = 42, class_weight = "balanced"
Inputs
Flow_Rate_LPM, Avg_Pressure_PSI, Occupancy_Status, Hour, Is_Weekend
Classes
no_leak, fixture_leak, valve_failure, mainline_break
Output
Diagnostic class
Loss estimation

LinearRegression

Input
The same diagnostic feature frame
Output
Estimated Leak_Loss_LPM

Training rows can be generated from baseline telemetry with the labels no_leak, fixture_leak, valve_failure and mainline_break. Synthetic leak rows use controlled changes in flow, pressure and loss rate, and include controlled noise.

  1. Baseline telemetry
  2. Flow and pressure changes
  3. Noise
  4. Labeled synthetic row
WHY?

The synthetic data supports controlled development and testing. It does not replace real labeled building telemetry.

Why context matters

High Activity

Event / High Activity context.
No anomaly detected.

HydroSentinel High Activity result with no anomaly detected
High Activity + Leak

Event / High Activity context.
Review required.

HydroSentinel High Activity plus Leak result requiring review
WHY?

Explainability baselines are created from synthetic training rows labeled no_leak. For each Occupancy_Status, the diagnostic artifact stores the median Flow_Rate_LPM and median Avg_Pressure_PSI. Overall median flow and pressure are also available as fallback values. When a pattern is flagged, the current telemetry is compared with the learned no leak baseline for that context.

Hidden Leak
HydroSentinel AI reasoning panel for the corrected Hidden Leak example
Flow23.1 L/min current flow

15.3 L/min learned baseline

51.0% difference
Pressure46.3 PSI current pressure

50.2 PSI learned baseline

7.8% drop
Outputsfixture_leak

99.6% classifier score

9.6 L/min estimated loss

The classifier score is not a calibrated physical leak probability.

Output

The API returns the diagnostic result, loss estimate, baseline comparisons, operating context, impact estimates, reasoning and telemetry used by the interface.

Telemetry chart
HydroSentinel telemetry chart showing flow and pressure over time

System

  1. Browser
  2. Next.js / React
  3. FastAPI
  4. Analysis service
  5. Validation
  6. Feature engineering
  7. RandomForestClassifier + LinearRegression
  8. Learned baseline comparison
  9. Impact calculation
  10. Structured API result
  11. Reasoning and charts
Frontend
Vercel
Backend
Render
Persistence
SQLAlchemy, Alembic, PostgreSQL targeted architecture

Reliability

  • Input validation.
  • Diagnostic model artifacts include a training fingerprint, schema version and learned baseline metadata.
  • Stale or incompatible artifacts are not reused.
  • If valid baseline metadata is unavailable, the UI reports that the comparison is unavailable.
  • The public demo uses a dedicated non persistent analysis endpoint.
  • Readiness checks database connectivity and the seeded scenario set.
  • The public endpoint includes rate limiting.

Why the baseline fix mattered

During final QA, the leak scenarios returned Review required but the explanation displayed 0 percent flow increase and 0 percent pressure drop.

WHY?

The diagnostic artifact did not contain the learned baseline metadata. The explanation code therefore used the current values as fallback baseline values, which produced a zero difference.

The fix added baseline metadata calculated from no_leak training rows. Baselines are stored by occupancy context. Diagnostic artifacts now include a schema version so stale artifacts can be rejected. A safe fallback and regression tests were also added.

Current limits

  • Synthetic telemetry.
  • No real building validation.
  • The classifier score is not a calibrated physical leak probability.
  • Estimated loss is a model estimate.
  • Human review is required.

Next technical work

  • Real telemetry.
  • Verified leak and no leak events.
  • Held out evaluation.
  • Precision and recall.
  • False positive and false negative analysis.
  • Score calibration.