How accurate is camera-based vital sign measurement?
In controlled testing against recognised clinical reference methods, the technology integrated within Vitals AI demonstrated strong performance. Results include heart rate within 1-3 bpm of hospital ECG, HRV with an SDNN RMSE of 6.3 ms against ECG-derived HRV, breathing rate within approximately 2.3 breaths per minute of a clinical reference, and approximately 95% of blood-pressure estimates within 10 mmHg of a cuff reference.
Vitals AI is built on independently benchmarked, dual-stream camera-based physiology technology. It combines remote photoplethysmography (rPPG) blood-flow signals with remote ballistocardiography (rBCG) heartbeat-related micro-movement signals, supported by real-time quality assessment and adaptive signal selection.
Independent benchmarking included Fitzpatrick skin types I-VI, diverse demographic groups, multiple age ranges and participants recruited across Australia, Poland, South Africa, Canada, the United States, China and Kenya. Because camera-based insights can be affected by skin tone, lighting, device quality and movement, Vitals AI applies quality controls to select the strongest available signal pathway.
This page brings together the research and validation behind Vitals AI. For the underlying mechanism, see how rPPG works; for the full marker list, see what it measures.
bpm vs hospital ECG
ms HRV SDNN RMSE
breaths/min
BP estimates within 10 mmHg
Independent benchmarking of the model used by Vitals AI.
Is rPPG accurate as a method?
Remote photoplethysmography has been studied for more than 15 years. A foundational 2008 study by Verkruysse, Svaasand and Nelson showed that ambient light and an ordinary camera could recover a blood-volume pulse signal from facial video (Verkruysse et al., 2008). Three years later, a 2011 study showed that the same principle could recover heart rate, breathing rate and heart rate variability from ordinary webcam footage, extending camera-based measurement beyond a single vital sign (Poh et al., 2011).
The field has expanded substantially since. A 2026 systematic roadmap screened more than 500 records down to 151 primary rPPG studies on the science and its path toward clinical use (Elgendi et al., 2026), and a 2025 audit analysed 100 studies built on public rPPG datasets to examine how representative the underlying data is (Bondarenko et al., 2025). Together, these reviews show that rPPG has developed into a substantial research field, with evidence spanning algorithms, datasets, real-world conditions and potential clinical applications.
The independent evidence is strongest for heart rate. A 2022 systematic review and meta-analysis of consumer-camera vital sign monitors found strong agreement with medical reference devices for heart rate, while noting there were too few studies to pool equivalent conclusions for blood pressure or breathing rate (Pham et al., 2022). A 2023 clinical meta-analysis pooling comparisons against ECG found a pooled bias of roughly -0.13 bpm, close to negligible at a population level, while still calling for broader clinical validation across settings (Bautista et al., 2023). A more recent 2026 comparison of webcam rPPG against ECG in 77 people reached a similar conclusion: average heart rate held up well, while individual-level measures like HRV were more limited (Woelk et al., 2026).
The optical principle is well established. Blood-volume changes with each heartbeat alter the amount of light absorbed and reflected by the skin. Traditional PPG captures that change with a sensor touching the skin, the same technology inside a pulse oximeter. rPPG detects the same underlying pulse signal remotely, from ordinary camera video, without contact.
This research establishes rPPG as a credible measurement method with a genuine, growing evidence base. It does not validate every commercial implementation: algorithm choice, signal processing and training data all affect how well a specific product performs, which is why the model used by Vitals AI is evaluated separately, through independent benchmarking rather than inferred from the field literature.
How is the Vitals AI model validated?
Vitals AI uses a dual-signal approach, combining remote photoplethysmography (rPPG) with remote ballistocardiography (rBCG). rPPG detects subtle changes in light reflected from the skin caused by blood flow, similar to the pulse signal a contact PPG sensor captures and undetectable to the human eye. rBCG adds a complementary signal from the minute body movements associated with each heartbeat. Analysing both together gives the model a richer physiological data set than either signal alone. This dual-signal model has been independently benchmarked against clinical reference devices.
The benchmarking compared model output against hospital ECG for heart rate and heart rate variability, a breathing reference device for breathing rate, and a cuff-based sphygmomanometer for blood pressure. It included participants recruited in Australia, Poland, South Africa, Canada, the United States, China, and Kenya, with Fitzpatrick I-VI skin types represented across the data rather than concentrated in one group. This is separate work from the field literature in the previous section: it is benchmarking of one specific model, run by an independent party, not a peer-reviewed study and not an Upvio-run trial.
These figures come from independent benchmarking of the model used by Vitals AI: heart rate within 1 to 3 bpm, HRV SDNN RMSE of 6.3 ms, breathing rate within about 2.3 breaths per minute, and roughly 95% of blood pressure estimates within 10 mmHg of a cuff reading. The full validation report is not yet public; see Sources and references for the independent field literature these results are benchmarked against.
HOW A READING BECOMES A VALIDATED RESULT
Camera video feed
+
rPPG / rBCG processing
=
Model output
→
Clinical refernce device
Simplified signal path: a facial video feed is processed into physiological signals, converted into a model output, then compared against a clinical reference device for validation.
How accurate is each vital sign, and what are the limits?
Not every output works the same way, and Vitals AI labels them accordingly. Some outputs are measured directly from the recovered physiological signal. Others are estimated from features within that signal, and one is modelled from a combination of signals rather than read from a single physiological marker.
Heart rate: measured
Heart rate has the deepest independent evidence base of any output on this page. It is calculated from the pulse waveform recovered from the camera signal, the same waveform Verkruysse and colleagues first demonstrated in 2008 (Verkruysse et al., 2008). Independent meta-analyses since have found strong agreement between camera-based heart rate and clinical reference devices, with one pooling a bias of roughly -0.13 bpm against ECG (Bautista et al., 2023; Pham et al., 2022). In independent benchmarking, the model used by Vitals AI measured heart rate within 1 to 3 bpm of hospital ECG, consistent with that wider evidence though not drawn from it.
Heart rate variability: measured
SDNN is calculated from the timing between successive beats in the recovered pulse signal, which places greater demands on signal quality than an average heart rate does: a missed or misplaced beat barely shifts a minute-long average but can meaningfully distort a beat-to-beat interval. The wider literature reflects that. A 2026 comparison of webcam rPPG against ECG in 77 people found HRV could be captured well at the group level, while individual-level estimates were more limited than average heart rate (Woelk et al., 2026). The benchmarked RMSE against ECG-derived HRV for the model used by Vitals AI was 6.3 ms, a figure that should be read as HRV-specific rather than assumed to match the heart rate result above.
Breathing rate: measured
Breathing rate is the second most studied output in the camera-based vitals literature after heart rate, though its evidence base is smaller (Selvaraju et al., 2022). Early webcam research already showed breathing rate could be recovered alongside heart rate and HRV from the same video feed (Poh et al., 2011). In independent benchmarking, the model used by Vitals AI measured breathing rate within about 2.3 breaths per minute of the clinical reference.
Stress: modelled
The stress score is calculated from HRV and other cardiovascular signals rather than measured from a single physiological marker. It shows patterns and changes over time rather than diagnosing a clinical condition, and should be read as a wellness indicator rather than a physiological measurement in its own right.
How accurate is camera-based blood pressure estimation?
Blood pressure is more challenging than heart rate, and the research base reflects that. Systematic reviews have found far fewer camera-based blood pressure validation studies than heart rate studies: a 2022 meta-analysis found too few eligible BP studies to pool into a single field-wide accuracy figure (Pham et al., 2022), and a 2022 review of 104 camera-vitals studies found blood pressure evidence much thinner than heart rate or breathing rate (Selvaraju et al., 2022). BP results should be read model by model rather than as a settled field-wide accuracy level.
Recent independent academic work shows genuine promise. A 2024 proof-of-concept study estimated blood pressure from facial video with encouraging early results (Trirongjitmoah et al., 2024), and a 2025 study reported that its facial-video BP model met AAMI and BHS error thresholds within its own test dataset (Yang et al., 2025). Those results show what individual research models have achieved in their own test populations. They do not mean camera-based blood pressure has been validated as a category to the same standard as a clinical cuff. ISO 81060-2, the standard commonly used for intermittent automated blood pressure monitors, specifically applies to devices that use a cuff.
Vitals AI treats blood pressure as an estimate rather than a direct measurement, calculated from features in the camera-derived signal and independently benchmarked against cuff readings. Roughly 95% of estimates in that benchmarking were within 10 mmHg of the reference reading. That makes the result useful for a wellness estimate and for tracking patterns over time. On the general wellness product it should not replace a cuff when a clinical reading is required, and is not described as clinical-grade. Blood pressure does sit within the CE MDR Class IIa certified scope of the separate regulated model (see Regulatory status).
It also helps to know why a camera estimate and a cuff rarely match to the millimetre. Blood pressure is not constant: it can move by around 10 mmHg within a few seconds, so two readings taken moments apart differ even on the same device. A cuff captures a single moment, while the camera estimate is a continuous average across the roughly 60 seconds of the scan. Certified cuff monitors also carry their own error, usually a standard deviation of about 3 to 5 mmHg. A gap between the two is therefore partly a difference in method and timing, and partly the cuff's own variability, rather than simply camera error. The estimate is still best read as a wellness figure for tracking patterns over time, not a replacement for a cuff when a clinical reading is needed.
What affects rPPG accuracy?
Several real-world factors influence how well a camera can recover a physiological signal:
Lighting: very low, uneven or rapidly changing light can make the pulse signal harder to recover.
Movement: significant head or body movement during a scan introduces noise and can reduce signal quality.
Camera quality: low-resolution or heavily compressed video can limit the physiological information available to the model.
Face visibility: the camera needs a clear view of the face during the scan.
Individual variation: optical physiological measurement can perform differently across people, which is why representative validation across skin tones and populations is important.
Independent testing also shows that strong results on standard benchmark datasets do not always carry over to harder conditions. A 2025 study testing eight rPPG methods found that elevated heart rates significantly reduced performance for most of the methods tested, while low illumination had a smaller but still relevant effect. Several deep-learning models also struggled when tested on data outside the dataset they were trained on (Acharya et al., 2025). A 2026 roadmap of the field reached a related conclusion: much of the published literature demonstrates technical feasibility, while broader clinical utility still needs to be established across real-world settings (Elgendi et al., 2026).
The practical lesson is that validation should resemble real deployment conditions, rather than only a controlled lab setting. The figures on this page describe performance under the conditions used in benchmarking, not a guaranteed result for every scan. Vitals AI is a general wellness tool and is not intended to diagnose or treat a medical condition.
Regulatory status
Vitals AI takes a best-of-breed approach to physiological modelling, combining specialist models within the Upvio platform. The model used by Vitals AI has CE MDR Class IIa certification and ANVISA registration:
CE MDR Class IIa under EU 2017/745, covering blood pressure, heart rate, heart rate variability and breathing rate.
ANVISA registration in Brazil, covering heart rate, heart rate variability, blood pressure and breathing rate.
These regulatory credentials attach to the model itself, not to Upvio as a company, and they sit alongside the independent benchmarking above: the model is both measured against clinical reference devices and formally assessed for regulated use. Vitals AI is offered as a general wellness tool by default. The regulated model is available as additional assurance for deployments that require a certified medical-device pathway, not a separate product most teams need to choose between.
Vitals AI aligns with the FDA's General Wellness Policy for low-risk devices and the Clinical Decision Support (CDS) carve-out (updated January 2026).
Does rPPG work on all skin tones?
Skin tone is an important validation variable in camera-based physiological measurement, because the signal the camera recovers depends on how light interacts with skin. Independent rPPG studies have found performance differences across skin tones in some methods, with the size of the effect varying by algorithm rather than being a universal failure of the technique (Nowara et al., 2020). A separate evaluation of demographic bias in rPPG reached a similar conclusion: some algorithms show a meaningful gap across skin tones, others do not, and the difference comes down to how the algorithm was designed and trained (Dasari et al., 2021).
The underlying datasets carry weight too. A 2025 audit of 100 studies built on public rPPG datasets found that many widely used datasets underrepresent darker skin tones, limiting how confidently a model trained or tested on them can be expected to generalise (Bondarenko et al., 2025). The lesson across all three: performance across skin tones should be measured, not assumed.
The model used by Vitals AI was evaluated across Fitzpatrick skin types I-VI, including a dedicated South African cohort added to strengthen representation of darker skin tones, alongside participants recruited in Australia, Poland, Canada, the United States, China and Kenya. Bias mitigation was part of the model’s training and evaluation process, and skin tone information used for this work was kept separate from participant identity throughout. The important point is that performance across skin tones was tested rather than assumed: the independent benchmarking covered Fitzpatrick I-VI and included participants recruited across seven countries.
This adaptive approach improves resilience where lighting, device quality, user movement or skin pigmentation may affect a single-method solution, supporting a more reliable and inclusive experience across diverse populations.
How is rPPG benchmarking actually done?
Validating a camera-based vitals model means comparing its outputs with established reference devices under defined test conditions, then reporting the error against those references rather than against another camera-based estimate. A benchmark is only as strong as its reference device, which is why hospital-grade equipment, not a consumer wearable, is the standard reference point for rigorous rPPG validation.
For the model behind Vitals AI, that meant:
Hospital ECG as the reference for heart rate and heart rate variability.
A clinical breathing reference device for breathing rate.
A cuff-based sphygmomanometer for blood pressure.
Reference device choice is only one part of the picture. Algorithm design also shapes how much noise a model has to contend with before it reaches the reference comparison. Early rPPG work relied on simple colour-channel averaging, sensitive to motion and lighting change. Later approaches, including CHROM, directly targeted sensitivity to motion by combining colour channels to cancel out specular reflection and movement artefacts (de Haan and Jeanne, 2013), and the POS approach set out broader algorithmic principles for isolating the pulse component of the video signal from noise (Wang et al., 2017). Those principles explain why the choice of signal-processing method affects real-world accuracy as much as the quality of the training data.
The model used by Vitals AI combines rPPG with rBCG and applies signal-processing methods to reduce noise from movement, lighting and the camera itself, before the benchmarking comparisons above take place. The technical signal pipeline is covered on the how rPPG works page; this page focuses on how the outputs were validated.
How should you evaluate an rPPG accuracy claim?
Accuracy figures are only useful if you know how they were produced. Three things are worth checking whenever a vendor or a study reports a number: the reference device used for comparison, whether the benchmarking was carried out independently, and the people included in the validation population.
Clinical reference devices such as hospital ECG and validated cuff-based blood pressure monitors provide a clearer basis for comparison than another estimated signal. Independent benchmarking adds evidence beyond a vendor's own internal testing. Validation across different countries and skin tones also gives a better picture of how a model performs outside a narrow sample, because results from one demographic group or recording environment do not automatically transfer to another (Dasari et al., 2021; Bondarenko et al., 2025).
Check what is being claimed and for which output. A strong heart rate result says nothing about how the same product estimates blood pressure or heart rate variability, since those outputs draw on different parts of the signal and carry different amounts of independent evidence (Pham et al., 2022; Selvaraju et al., 2022). A broad claim, “our vitals are clinically accurate”, without naming a reference device or metric, is harder to verify than one that states a number, a comparison device and a population.
Those are the same standards used for the figures reported on this page: independent benchmarking against clinical reference devices, multi-country data, and Fitzpatrick I-VI coverage.
Sources and references
These sources cover the independent field literature behind rPPG as a method — organised below by foundational research, clinical reviews, and fairness and methodology work. The accuracy figures reported for the model used by Vitals AI come from separate, independent benchmarking, covered in model validation above.
Independent rPPG literature (4)
01
Verkruysse, Svaasand & Nelson (2008)
Optics Express, 16(26), 21434-21445.
The foundational study demonstrating blood-volume pulse extraction from facial video using ambient light.
02
Poh, McDuff & Picard (2011)
IEEE Transactions on Biomedical Engineering, 58(1), 7-11.
Early demonstration of multi-signal extraction from consumer webcam video.
03
Wang, den Brinker, Stuijk & de Haan (2017)
IEEE Transactions on Biomedical Engineering, 64(7), 1479-1491.
The optical and physiological principles behind rPPG algorithms.
04
de Haan & Jeanne (2013)
IEEE Transactions on Biomedical Engineering, 60(10), 2878-2886.
The CHROM method, a landmark algorithm for motion-robust pulse extraction.
Clinical and systematic reviews (5)
01
Pham et al. (2022)
Journal of Clinical Monitoring and Computing, 36(1), 41-54.
Systematic review: strongest evidence for heart rate, thinner evidence for blood pressure and breathing rate.
02
Bautista et al. (2023)
Journal of Clinical and Translational Science, 7, e129.
Clinical heart-rate meta-analysis comparing contactless measurements against ECG.
03
Elgendi et al. (2026)
npj Digital Medicine.
Current field roadmap covering 151 primary rPPG studies and the path toward clinical use.
04
Selvaraju et al. (2022)
Sensors, 22(11), 4097.
Reviews the breadth of camera-vitals research; breathing rate is well represented, blood pressure much thinner.
05
Trirongjitmoah et al. (2024)
Heliyon, 10(5), e27113.
Independent facial-video blood-pressure proof-of-concept.
Fairness, conditions and methodology (6)
01
Dasari et al. (2021)
npj Digital Medicine, 4, 91.
Shows that demographic effects vary by algorithm and training design.
02
Bondarenko, Menon & Elgendi (2025)
npj Digital Medicine, 8, 593.
Examines representation of darker skin tones in public rPPG datasets.
03
Nowara, McDuff & Veeraraghavan (2020)
CVPR Workshops, 284-285.
Finds skin-tone effects are algorithm-dependent rather than universal.
04
Acharya, Saakyan, Hammer & Drimalla (2025)
npj Digital Medicine, 8, 744.
Reports real-world limits involving low light, elevated heart rates and cross-dataset generalisation.
05
Woelk, Garfinkel, Mayiwar et al. (2026)
Behavior Research Methods, 58, 135.
Finds HRV is more reliable at the group level than the individual level.
06
Yang, Gu, Liu et al. (2025)
Physical and Engineering Sciences in Medicine, 48(4), 2059-2067.
Independent facial-video blood-pressure study reporting thresholds within its own dataset.
Questions, answered
Is rPPG accurate?
Yes, for heart rate in particular. The field has more than 15 years of published research and over 100 peer-reviewed studies (Verkruysse et al., 2008; Elgendi et al., 2026). Product-level accuracy still depends on the specific model: the model used by Vitals AI has been independently benchmarked against hospital ECG, measuring heart rate within 1 to 3 bpm.
How accurate is the blood pressure estimate?
Blood pressure is treated as an estimate, not a direct measurement. In independent benchmarking, roughly 95% of the model's estimates fell within 10 mmHg of a cuff reading, useful for a wellness estimate and for tracking patterns over time, but not a substitute for a cuff when a clinical reading is required.
Does it work on all skin tones?
The model used by Vitals AI was evaluated across Fitzpatrick skin types I-VI, including a dedicated South African cohort, alongside participants recruited in Australia, Poland, Canada, the United States, China and Kenya. Independent research shows skin tone effects in rPPG are algorithm-dependent rather than universal, which is why performance across skin tones was measured rather than assumed.
Is this a medical device?
The general wellness product is not a medical device: it does not diagnose, treat or monitor a medical condition, and its results should not replace a clinical measurement or a qualified healthcare professional's judgement. Separately, a CE MDR Class IIa certified model is available for regulated use.
How was it validated?
The model used by Vitals AI was independently benchmarked against clinical reference devices: hospital ECG for heart rate and heart rate variability, a clinical breathing reference for breathing rate, and a cuff-based sphygmomanometer for blood pressure, across seven countries and Fitzpatrick I-VI skin types.
Does rPPG work in low light?
Low light is one of the harder conditions for camera-based measurement, because the model needs enough reflected light to detect the pulse signal. Independent research has found accuracy can decline in low illumination and at elevated heart rates (Acharya et al., 2025). Vitals AI expects reasonably lit, front-facing conditions, and scan guidance reflects that.
What age range is Vitals AI recommended for?
Vitals AI is recommended for people aged 12 to 70. Outside that range, camera-based measurement is less well established, so results should be treated with more caution.
How does rPPG compare to a pulse oximeter?
A pulse oximeter uses contact PPG, with a sensor touching the skin to measure optical changes associated with blood flow and oxygen saturation. rPPG applies the same underlying optical principle remotely, recovering a pulse signal from camera video rather than a contact sensor. Vitals AI does not measure oxygen saturation. Its outputs include heart rate, heart rate variability, breathing rate, an estimated blood pressure reading and a wider set of wellness indicators.
Why doesn't the camera match my cuff exactly?
Blood pressure changes constantly, by around 10 mmHg within a few seconds, so no two readings line up perfectly, even on the same device. A cuff takes a single-moment snapshot, while the camera estimate is an average across the roughly 60-second scan, and certified cuffs carry their own error of about 3 to 5 mmHg. Expect the two to sit in a similar range rather than match to the millimetre, and use the camera figure to track patterns over time.
Related reading and next step
Go deeper, then explore the product.
Try Vitals AI
Run a scan or talk to the team about using it in your workflow.