Two large randomized experiments, published within a year of each other, are the closest thing this field has to a clean test, and both landed in the same place: real gains in participation and some behaviors, but no measurable effect on the money outcomes.[1][3]
The Illinois Workplace Wellness Study
The Illinois Workplace Wellness Study randomly assigned about 4,800 university employees either to be offered a two-year wellness program or not, then tracked screenings, biometrics, medical claims, and job outcomes.[1] Participation and health screening rose and stayed higher, and more employees in the treated group reported having a primary care physician after two years.[2] The program produced no significant causal effect on total medical spending, other health behaviors, employee productivity, or self-reported health after more than two years.[1] The experiment also exposed why observational studies look so much rosier: in the year before the program began, the employees who chose to participate already had lower medical costs and healthier behaviors than those who did not, and the trial's confidence intervals ruled out 84% of the earlier published estimates for savings on medical spending and absenteeism.[1][2]
The BJ's Wholesale Club trial
Song and Baicker's cluster-randomized trial assigned 160 BJ's Wholesale Club worksites, covering 32,974 employees, either to receive a multi-module wellness program or to serve as controls, then followed them for 18 months.[3] Workers at the treated sites were about 8 percentage points more likely to report exercising regularly and about 14 points more likely to say they were actively managing their weight.[3] Those self-reported behavior gains did not carry through to anything measured from records: cholesterol, blood pressure, glucose, and body-mass index were unchanged, and there was no significant difference in medical spending, health-care use, absenteeism, job tenure, or performance-review scores.[3] Because the study randomized whole worksites and drew on clinical and claims data rather than volunteers, the null results are hard to explain away as a measurement artifact.