Machine Learning - Solving Yet Another Problem? Perhaps less need for arterial access?
Matthew J Rowland, MD, Ethan Sanford MD, Shawn Jackson MD PhD
There is no question that end-tidal carbon dioxide (EtCO2) monitoring has changed the practice of anesthesia. Outside of being the gold standard for confirming endotracheal intubation, end-tidal CO2 is a critically valuable metric for the assessment of ventilation and cardiac output. However, while we often use EtCO2 as a surrogate for the arterial partial pressure of carbon dioxide (PaCO2), they are not the same. In children with normal cardiovascular anatomy and healthy lungs, it can often be assumed that the PaCO2 is 2-5 mmHg higher than the EtCO2, previous studies have highlighted discrepancies between EtCO2 and PaCO2.1 This is especially true in children with lung disease.1 Dead space ventilation (areas of lung ventilation without perfusion) is likely the cause in most cases.
In anesthesia, both underlying lung disease and changes in peri-procedural physiology make estimation of PaCO2 based on end-tidal more precarious. Examples include cardiopulmonary bypass surgery, neurosurgery, thoracic surgery and cases requiring significant fluid resuscitation. Thus, monitoring of PaCO2 levels is one common indication to place arterial access in patients.
Yet, the future of the arterial line is becoming less certain. As we heard during the most recent Society of Pediatric Anesthesia meeting in Denver from Dr. Michael Fiedorek, numerous new devices are becoming available that reliably estimate beat by beat arterial blood pressure accurately, including a device approved in neonates (Boppli). While these devices are promising for monitoring continuous blood pressure, they do not provide the same laboratory information as we are accustomed to obtaining with an arterial line. However, perhaps AI can provide a solution here.
Today, we review a recent study from Ju Park and colleagues that uses machine learning to address a specific need: bridge the gap between EtCO2 and PaCO2 and, thus, create novel non-invasive monitoring techniques.
Original Article
Park JH, Cho C, Kim HS, et al. Development of an Arterial Carbon Dioxide Estimation Model Using End-Tidal Carbon Dioxide Levels during Surgery in the Pediatric Population. Anesthesiology. Published online June 17, 2026. doi:10.1097/ALN.0000000000006207
The authors looked at over 8,000 pairings between EtCO2 and PaCO2 in 3,500+ pediatric patients in a large pediatric database (VitalDB). They used 70% of this data to train four different machine learning algorithms. Then they internally validated their model on another 15% of the database and tested it out on the last 15% of data in database. Lastly, they externally validated the model on two different external data sets.2 The analysis included any patient less than 19 years old undergoing general anesthesia with an endotracheal tube and arterial line. Including ~35% of patients that were ASA PS III. Notable exclusions from the data include any cardiopulmonary bypass cases, one-lung ventilation, laparoscopic and thoracoscopic cases.2
The authors made several adjustments to limit bias, including avoiding blood gases at the start or end of cases and allowing for lags in data between drawing the sample and determination of the PaCO2. The authors also split the models into two distinct groups - less than 6 years old and 6 years old and above.2
The four models were indeed able to learn and predict PaCO2 fairly accurately. The best model (Gradient Boosting) had a mean error of 2.73 mmHg in predicting the PaCO2 with a percentage error of 6.75%.2 While this is impressive, it does translate to a potential alteration in parameters like cerebral blood flow by 6-12%, a relatively large difference. Figure 5 from the study is a Bland-Altman plot demonstrating less difference from actual values predicted by the model (blue dots) than for the end-tidal alone model (red dots).
Many (us included) gloss over when trying to understand and interpret studies of machine learning models, but the devil is certainly in the details. Very simply put, the models utilize all of the data made available (surgical, hemodynamic, anesthetic, ventilation, laboratory, and patient) to develop predictions. The model is essentially a new monitor combining all the variables we look at daily to make a prediction, much as we do. Interpreting the results is highly analogous to standard clinical research critiques i.e. what was the population, do the variables used to predict make sense or feasible to extract, do the results yield meaningful clinical information or is it just window dressing. This study excluded the patients more likely to have difference between end-tidal and PaCO2. This strengthens the model for prediction in healthy kids but precludes use in patients we are more concerned about. Similarly, as dead space ventilation increases, this model is likely to be less accurate. Lastly, the error rate, while close, is not yet close enough when more accurate information can be obtained via an arterial catheter. If making clinical decision regarding cerebral blood flow and cerebral perfusion pressure needs, is it better to have an estimate or the real data? Finally, the intelligence gained from models must be significantly different from what we already know and be modifiable.
However, this technology does give us hope for a future where machine learning may allow for a model that is even more accurate and precise. Additionally, even in current state, one could consider applying this technology to estimate PaCO2 in patients where arterial line access is unobtainable or challenging, particularly in institutions where transcutaneous CO2 monitoring (tcPCO₂) is less available. Threshold alarms could be set to help us recognize and act when deviations occur.
Yet, what we do with new information provided by machine learning is also important. The hypotension prediction index trial of a machine learning algorithm which predicted hypotension failed to change time with hypotension because clinicians either didn’t act on the alarm (ie they thought the alarm was irrelevant) or there wasn’t enough time between predicted hypotension and actual hypotension to create an opportunity to change management. Alternatively, the HYPE trial lowered the threshold for intervention based on the machine learning model and resulted lower amounts of hypotension.3 In other words, if and how we can act on these models matters and even then, it’s not completely clear if action will change outcomes.
What are your thoughts on the machine learning and its ability to predict arterial carbon dioxide levels? Is this the start of the end of the arterial lines? Send your thoughts to Myron (myasterster@gmail.com and he will post in a Friday reader response..
References
1. Yang JT, Erickson SL, Killien EY, Mills B, Lele AV, Vavilala MS. Agreement Between Arterial Carbon Dioxide Levels With End-Tidal Carbon Dioxide Levels and Associated Factors in Children Hospitalized With Traumatic Brain Injury. JAMA Netw Open. 2019;2(8):e199448. Published 2019 Aug 2. doi:10.1001/jamanetworkopen.2019.9448
2. Park JH, Cho C, Kim HS, et al. Development of an Arterial Carbon Dioxide Estimation Model Using End-Tidal Carbon Dioxide Levels during Surgery in the Pediatric Population. Anesthesiology. Published online June 17, 2026. doi:10.1097/ALN.0000000000006207
3. Wijnberge M, Geerts BF, Hol L, et al. Effect of a Machine Learning-Derived Early Warning System for Intraoperative Hypotension vs Standard Care on Depth and Duration of Intraoperative Hypotension During Elective Noncardiac Surgery: The HYPE Randomized Clinical Trial. JAMA. 2020;323(11):1052-1060. doi:10.1001/jama.2020.0592



The HPI paragraph is the part I'd make people read twice. A model that predicts correctly and changes nothing is an expensive way to be right. What strikes me about building this on end-tidal CO2 in particular is how that monitor earned its place: in the mid-1980s it answered one question immediately and unambiguously, which is whether the tube is in the trachea, and unrecognized esophageal intubation essentially stopped being a way people died. That is a monitor with an unmistakable action attached to its output. An estimated PaCO2 carrying a 6.75% error, in exactly the population you excluded from the training set, is a different kind of object. Your closing criterion is the right one, and I'd add that it has to arrive with the action already attached.