Blog
Part 1 – From Physics to Machine Learning: Building a System-Level, ML-Based BEV Digital Twin
Overview
Digital twins have moved from pilot projects to everyday engineering tools. Across industries, teams are under pressure to deliver more complex products on shorter schedules, and physical testing alone can’t keep up. A digital twin predicts how a system will behave, such as its performance, wear, or risk of failure, before hardware exists and after it is in the field. The business case is showing up in the numbers. McKinsey has reported that digital twins have cut total development time by 20 to 50 percent for some users, and that products developed as twins can reach production with 25 percent fewer quality issues. The market is following: The Business Research Company estimates it will grow from about $29.6 billion in 2025 to $42 billion in 2026. The value of a twin also depends on where it runs. A high-fidelity physics model suits design work, but it is often too heavy for the constrained hardware found in the field. This series shows how to close that gap.
Two Questions Every Digital Twin Project Eventually Asks
Every digital twin conversation eventually runs into the same two forks in the road. The first is a runtime question: does the twin need to run online, in step with a real system, or is it purely an offline analysis tool? The second is a location question: does it live in the field, on the constrained hardware of a vehicle’s own controller, or on the cloud, where compute is cheap and elastic? The answers determine everything downstream, from model fidelity to the programming language you end up writing.

This two-part series walks through both ends of that spectrum using a battery electric vehicle (BEV) digital twin built in GT-SUITE. Part 1 (this post) shows how a high-fidelity offline model gets distilled into fast machine learning metamodels, and briefly covers the on-cloud alternative with GT-Play. Part 2 deploys those metamodels onto real embedded hardware, an ESP32 microcontroller, as a working in-field digital twin.
The Starting Point: A High-Fidelity Offline BEV Model in GT-SUITE
The starting point is an integrated, multi-domain BEV model combining three physics-based sub-models:
● 1D Thermal: vehicle thermal management (VTM) network
● P2D Battery: a pseudo two-dimensional electrochemical battery model resolving solid and electrolyte phase diffusion and charge transfer kinetics
● 1D Vehicle: longitudinal vehicle dynamics (aero, rolling resistance, road grade)

This is the model of choice for design exploration. It is physically transparent and highly accurate, and it forms the ground against which every faster model is validated. GT-AutoLion’s approach to physics-based battery modeling, described in this companion post on modeling physical digital twins, covers the electrochemical side of this picture in more depth.
Making It Faster: Runtime Improvement Techniques
Before jumping to machine learning, GT-SUITE offers several routes to speed up the physics itself:
Thermal: physical reduced order models (ROMs) and VTM-to-FRM (Fast Running Model) converters
Battery: single particle models with real-time deployable anode and cathode bulk-surface splits
Vehicle and engine: map-based approaches or GT-POWER-xRT

These techniques already deliver strong real time factors (see the bar chart above, where the orange “XiL Compatible” bars sit far below the blue baseline across every drive cycle).
The On-Cloud Digital Twin Alternative: GT-Play
Not every digital twin needs to live on embedded hardware. For applications where the model can be hosted centrally and accessed remotely, GT-Play offers an on-cloud path: model developers push new models and parameter changes from a traditional GT-SUITE desktop installation to a web hosted server, which in turn dispatches jobs to an HPC scheduler and returns results to any connected device through a lightweight portal.

In-Field Digital Twins Need Light, Fast Models
Embedded microcontrollers used in battery management systems (BMS) and ECUs are built around two very different memory pools: flash, which stores the model, and RAM, which computes it. Both are compact by design, typically a few megabytes of flash and only a few kilobytes to a couple of megabytes of RAM. For a microcontroller class target, though, even a lightweight physical ROM can still ask for more flash and RAM than is available. That is the gap machine learning is built to close, to get an overview of GT’s ML capabilities you can visit Part 2 of GT’s machine learning blog series.

This is precisely the opportunity a light, fast machine learning metamodel is designed to capture: keep the accuracy of the physics based model, and shrink the memory footprint.
From Physics Model to Machine Learning Metamodel
The conversion follows a repeatable pipeline: start from the high-fidelity integrated model, run a structured design of experiments (DOE) to generate training data, preprocess, train, and postprocess that data into a metamodel, and validate first at the standalone subsystem level before assembling a full integrated system level ML model.

Generating the Training Data
The DOE swept 10 drive cycles (standard and real-world cycles derived using GT-RealDrive), 3 initial battery temperatures (10°C cold, 25°C room temperature, 45°C hot), and 7 initial states of charge: 210 training cases in total. Run natively, this would represent roughly 10 days of physical testing. Using distributed computing on the multiphysics model, the entire dataset was generated in about 6 hours of simulation.

Preprocessing: Teaching the Network About Time
Feature engineering is a vital aspect of machine learning and GT-MLA provides some essential plots like sensitivity analysis and correlation tables for identification of sensitive inputs and elimination of insensitive ones to improve training efforts alongside data filtering and sampling interval controls. Because the system is dynamic, the metamodel needs to know not just the current inputs/outputs but also recent history. GT-SUITE’s Machine Learning Architect (MLA) tool supports this through factor/response tap delay identification: partial autocorrelation (PACF) for response delays and cross correlation (CCF) for factor delays. The result feeds a NARX (Nonlinear AutoRegressive model with eXogenous inputs) architecture, where lagged inputs and lagged outputs are explicitly fed back into the network.

Training Workflow
In MLA, first, the available simulation/DOE dataset is divided into training, validation, and test data; then GT-SUITE creates all specified neural-network configurations and trains each candidate sequentially. Because neural-network training has stochastic variation, each configuration can be trained multiple times. After training, the models are compared using a metric such as validation RMSE, and the network with the best validation performance is selected as the final metamodel; the separate test dataset, which was not used during training, is then used to evaluate how well the selected model generalizes. More details can be found here.

The Payoff: 280 Times Faster, 1,000 Times Lighter
Once trained, the integrated ML model reproduces the physics based integrated BEV model’s outputs (battery temperature, motor temperature, HV/LV load, motor power, motor and battery heat generation, state of charge, battery current, and terminal voltage) while running roughly 280 times faster and about 1,000 times lighter than the source model (including solver files) it was trained on. To learn more about this you can read our publication: ML based fast & light Digital-Twin of an Integrated BEV (2025) – GT_IND_AM_VEHICLE_PUBLIC.pdf

What’s next? From ML to Silicon
The results are compelling: the integrated ML metamodel can reproduce the behavior of the high-fidelity BEV digital twin while running roughly 280 times faster and 1,000 times lighter than the source physics-based model. But for an in-field digital twin, speed and size improvements on a workstation are only part of the story. The next question is more practical: can the trained model run in real time on the resource-constrained hardware available in the field?
In Part 2, we’ll take the GT-generated ML metamodels from the desktop to an ESP32 microcontroller, using Simulink as the bridge between GT-SUITE and the embedded target. We’ll start with a single NARX model, validate its live embedded performance against the GT-SUITE reference, and then scale the approach toward a full closed-loop system of interconnected battery, motor, thermal, and vehicle models. Stay tuned for Part 2: From ML to Silicon – Deploying a System-Level, ML-Based BEV Digital Twin on a Microcontroller.
In the meantime, browse our blog page for additional insights on digital twins and ML/AI. Follow our LinkedIn channel to stay updated on the latest advancements in our solutions and subscribe to the GT blog to receive updates on new AI capabilities, engineering solutions, and industry trends.
Frequently Asked Questions
1. What is a digital twin?
A digital twin is a virtual model that represents the behavior of a physical system. It can use physics-based models, machine learning, or a combination of both to predict system performance, identify potential issues, and support engineering decisions before and after hardware is built.
2. Why use machine learning for a digital twin?
Machine learning can create lightweight metamodels that reproduce the behavior of high-fidelity physics-based models while requiring significantly less computational power. In this BEV example, the integrated ML model runs approximately 280 times faster and is 1,000 times lighter than the original physics-based model, making it more suitable for real-time and resource-constrained applications.
3. Can a machine learning digital twin run on embedded hardware?
Yes. A trained ML metamodel can be optimized for deployment on resource-constrained embedded hardware such as microcontrollers. In Part 2 of this series, the GT-generated ML metamodels are deployed on an ESP32 microcontroller and evaluated for real-time performance as part of an in-field BEV digital twin.