Skip to content

Operational AI for Data Centers

July 7, 2026

Data centers span a wide spectrum, from small lab environments to hyperscale campuses. Somewhere in between lie colocation facilities hosting multiple customers and enterprise data centers supporting business applications. Despite their differences in scale, they all share a common challenge: managing redundancy, capacity, and infrastructure dependencies across a fragmented landscape of systems and data.

This is where Operational AI transforms the equation. Willow’s rich ontology supports data center assets and relationships in the Knowledge Graph. Skills and Insights enable real-time monitoring and control. Building on this foundational data layer, facility teams can perform capacity management to optimize resource usage across power, cooling, and space.

Why Ontology Matters

Data centers are highly interconnected systems. IT racks depend on upstream power infrastructure. Every kilowatt consumed generates heat that must be dissipated by cooling systems. Failures in one part of the system ripple across others. 

By augmenting Willow’s ontology to include constructs like Data Halls to model spaces where IT Racks are housed, Busways for the electrical distribution, Automatic Transfer Switches (ATS) that automatically switch load between utility power and battery for Uninterruptible Power Supplies (UPS) and more, data centers can be represented in the Knowledge Graph. 

When a problem occurs, it becomes very easy to instantly know the physical location of a Rack. If a PDU has a voltage spike, it is clear which UPS to check. When a breaker trips, it is clear which IT Racks are impacted. This representation enables teams to navigate from a high-level view down to granular assets, while preserving context along the way. 

Modeling twins with relationships in Willow’s ontology helps define redundancy in power systems- be it N+1 or 2N. Utility power feeds into UPS systems that connect to ATS units routing power to PDUs, which supply IT racks.

These relationships allow Willow to represent redundancy configurations such as N+1 with one extra component for failover, and alternatively 2N fully redundant systems, where two completely independent UPS systems are each capable of supporting 100% of the critical load. As an example, if the load requires 4 UPS modules, the 2N design deploys 8 modules, 4 in Path A and 4 in Path B, with either path able to carry the entire load independently. If one path fails, the other maintains uptime. Willow captures this redundancy explicitly, allowing teams to understand risk exposure and failover paths.

With Willow’s ontology in place, organizations can leverage a consistent data model across all systems.

Unlocking Real-Time Monitoring and Operational Insights

While ontology provides foundational structure, additional value is realized when it is combined with live telemetry being monitored 24×7 with Skills to generate Insights. Willow can integrate signals from building management systems (BMS), IoT sensors, and Data Center Infrastructure Management (DCIM) platforms into a unified data layer, transforming the digital twin into a real-time operational model.  

Live telemetry from these systems is normalized and processed with Skills and Insights, enabling visibility across critical operational domains. Cooling can be monitored continuously through CRAC performance, airflow patterns, and temperature gradients between hot and cold aisles. Subtle shifts, such as a widening temperature differential, can signal airflow imbalance or containment issues before they escalate. At a more granular level, equipment telemetry such as fan amperage or discharge air temperature provides early indicators of wear, inefficiency, or misconfiguration. Anomalies can be traced across upstream and downstream dependencies, revealing how localized issues propagate through the system. 

Power monitoring allows teams to track real-time load at the rack PDU level, power distribution across circuits, and variations that might signal imbalance or overload risk. Willow can ingest electrical data from meters, panelboards and switchboards. Per-phase current data can expose imbalances that average values might obscure, while harmonic distortion and reactive power can point to inefficiencies or stress on electrical infrastructure.

Here we see the Excessive Switchboard Power Consumption Skill running on Service Switchboard A to continuously monitor power draw, generating an Insight when it exceeds adaptation baselines for normal operations. This is particularly useful for detection and diagnosis of transient events. 

When a spike is observed at a PDU, Willow can correlate it with upstream context, such as a recent ATS transfer or a fluctuation at the UPS, while simultaneously identifying downstream systems that may be affected. What would traditionally require manual investigation across multiple tools becomes a connected, explainable insight.  

Actioning Insights ultimately enables a shift from reactive monitoring to proactive operations. Teams can detect early deviations from expected behavior, identify inefficiencies such as uneven cooling or unbalanced electrical loads, and perform root cause analysis by traversing the relationships modeled in the system. A temperature anomaly in a rack can be traced to a specific CRAC unit, just as a power irregularity can be followed upstream through the PDU, busway, and UPS to its source. 

The result is real-time situational awareness, as organizations gain a unified view of operations, enriched with context. This reduces time to resolution and improves operational efficiency.

The Capacity Management Challenge

Capacity management in data centers is multi-dimensional, requiring teams to balance power, cooling, and physical space. For example, adding new racks increases power demand, which in turn generates higher heat loads. This additional thermal output must be absorbed by cooling systems, which may themselves be constrained by airflow design or equipment limits. Without coordination, incremental changes in one area can create bottlenecks in another. Inefficiencies such as stranded capacity can occur, where power is available but cooling is not, or overprovisioning, which can lead to wasted resources. 

Willow addresses these challenges by bringing all capacity dimensions into a unified analytical framework. This enables a holistic understanding of how capacity is allocated, consumed, and constrained across the environment.

Teams can analyze peak and average power loads, assess utilization across UPS systems and PDUs, identify remaining headroom, and determine where additional workloads can be deployed. This level of visibility helps maximize existing infrastructure while maintaining resilience and operational safety. Cooling capacity can be evaluated alongside power by looking at peak and average loads on CRAC units. Organizations can assess how changes in workload distribution impact thermal stability. Physical space is also incorporated into the model, allowing teams to optimize density and determine how space is being used in relation to both power and cooling availability. 

Willow Copilot introduces a natural language mechanism to interact with this data. Teams can prompt for racks with the highest electrical capacity, available headroom over current HVAC load, or rPDUs running over capacity, and more. Here’s an example:

By unifying data across power, cooling and space, Willow simplifies capacity management. Organizations can maximize infrastructure utilization, avoid unnecessary capital expenditures, and reduce the risk of outages caused by capacity constraints.

Conclusion

Willow delivers operational efficiency for data centers with a rich ontology, real-time telemetry monitoring with Skills and Insights, and capacity management, all in a unified platform. Teams benefit from structure, visibility and AI-driven analytics across power, cooling, and space. This integrated approach enables organizations to evolve from fragmented tools and reactive workflows toward a more predictive and optimized model of operations. With Willow Copilot, these capabilities become even more accessible, allowing teams to explore complex infrastructure data through natural language and quickly uncover actionable insights. Operational AI paves the path towards improving efficiency, resilience, and decision-making across the board.

Ready to transform?

See how Willow can help you cut costs, save energy, and operate your buildings more efficiently and sustainably.