Full methodology. Nathalia Hyland, Research & Editorial Lead.
Reader-facing summary: How We Counted Every Vote in the Summer 2026 issue.
The study measures which tools Direct Primary Care physicians in the My DPC Story audience report using, and how those users rate the tools they have chosen.
It does not measure comparative preference across tools. Respondents rated only the tool they currently use in each category. A physician using Spruce Health rated Spruce Health; they did not evaluate the alternatives. Category results therefore describe adoption and satisfaction among existing users, not head-to-head performance.
This distinction matters for interpretation. A tool with a small, well-matched user base can score highly without being the better choice for a physician outside that base.
There is no complete registry of Direct Primary Care physicians in the United States, so no probability sample of this population is possible. Any survey of DPC physicians is a convenience sample, and this one is no exception.
The sampling frame is the My DPC Story audience: newsletter subscribers, podcast listeners, and social media followers. This is a self-selected subset of an already self-selecting population, since physicians who operate DPC practices have opted out of conventional practice models.
Reported recruitment source (n = 131):
| Source | Responses |
|---|---|
| My DPC Story newsletter | 63 |
| My DPC Story podcast | 25 |
| Other | 30 |
| Social media | 8 |
| Another physician | 5 |
Peer referral was minimal. Recruitment included a raffle whose platform awards additional entries for sharing the survey link (section 3). Of 61 raffle entrants, one referral was credited. The survey's own referral-credit field drew five entries, all naming the same person. Both sources agree that response clustering through physician networks was not a material factor in this cycle.
Observed incentive effect. Responses arrived at 4.0 per day during the 20-day raffle window and 1.2 per day across the remaining 43 days. Eighty of 131 responses fall inside the raffle window. The incentive materially increased response volume.
Fifty-three percent of respondents completed the survey without entering the raffle, so the sample is not composed solely of incentive-motivated participants.
Practice size
| n | |
|---|---|
| Solo | 90 |
| 2 to 3 physicians | 33 |
| 4 to 6 physicians | 4 |
| 7 or more | 4 |
Years operating in DPC
| n | |
|---|---|
| Less than 1 year | 35 |
| 1 to 2 years | 25 |
| 3 to 5 years | 34 |
| 6 to 10 years | 30 |
| 10+ years | 7 |
Active members
| n | |
|---|---|
| Fewer than 100 | 37 |
| 100 to 300 | 38 |
| 301 to 600 | 27 |
| 600+ | 29 |
Geography: 40 states represented. California (22), Texas (10), Florida (7), North Carolina (7), South Carolina (6), Utah (5).
The sample skews strongly toward solo practices, which represent 69 percent of respondents. Findings should be read as most representative of solo and very small DPC practices.
Nine categories were assessed: patient communication, membership billing and payments, scheduling and access, telemedicine and virtual care, AI and automation (non-EHR), practice operations and admin, HR and payroll, labs and imaging, and patient education and engagement.
Each category presented a list of named tools plus an "Other (not listed)" option with a free-text field. Respondents who selected a tool then answered three rating items:
The labs and imaging category substituted reliability and fair pricing for ease of use and value, reflecting the different nature of a lab relationship.
An optional free-text field asked why the respondent recommends the tool. A closing item asked which single non-EHR tool the respondent would keep if limited to one.
Known instrument limitations
No importance weighting. Respondents rated satisfaction but were not asked how much each category matters to their practice. A low rating in a category central to daily work counts the same as a low rating in a peripheral one.
Category boundaries are not mutually exclusive. Several all-in-one platforms span multiple categories. A physician using an integrated system could reasonably report it under communication, billing, scheduling, and telemedicine, or report a category-specific alternative. Section 6 addresses the coding consequences.
Option lists were incomplete. See section 6.
"Other (not listed)" was the most-selected option in four of nine categories: patient communication (58), scheduling (51), telemedicine (53), and AI and automation (69 of 101 responses).
This means the presented option lists did not capture the tools actually in use. The free-text entries show why. All-in-one platforms including AtlasMD, SigmaMD, and Cerbo appeared repeatedly across categories with inconsistent spelling. In scheduling, a physician using an integrated platform could select "Native EHR tool" or type the platform name, and respondents did both, splitting one underlying answer across two codes.
Recoding audit. Free-text entries were normalized to canonical tool names using a documented rule set: case and punctuation normalization, then pattern matching against a fixed list of known platforms. Entries indicating no tool in use ("none", "n/a") were removed rather than counted as Other. Blank free-text entries under Other were dropped as unattributable.
Result: no category leader changed. All nine published winners hold under recoding.
Second and third positions do move:
| Category | Change after recoding |
|---|---|
| Patient communication | AtlasMD enters third (13) |
| Membership billing | SigmaMD enters third (10), displacing Stripe |
| Scheduling | AtlasMD enters third (10), displacing Google Calendar |
| Telemedicine | AtlasMD enters third (16), displacing Zoom |
| AI & automation | Open Evidence rises to second (16), displacing Freed |
Two cautions on these. AtlasMD at 16 in telemedicine sits close to the 20-20 tie between Google Meet and Spruce Health, a gap well inside sampling noise at this sample size. And Open Evidence appears both as a named option in patient education and as free text under AI and automation, indicating that respondents do not place it consistently.
A long tail of single-mention tools remains unmapped: 30 distinct entries in patient communication, 26 in scheduling, 24 in HR and payroll. These are genuine one-off tools rather than coding errors, and they indicate a fragmented market that no fixed option list would fully capture.
Rating completeness. Among respondents who selected a tool, both rating items were completed at high rates: 122 of 129 in patient communication, 125 of 130 in scheduling, 118 of 126 in labs and imaging. The lowest completion was AI and automation at 90 of 101.
Missing data. Ratings were analyzed pairwise. A respondent who selected a tool but skipped a rating item contributes to the adoption count for that category and not to the rating means.
Duplicate submissions. No formal screening was performed. The survey itself is anonymous and carries no identifiers, so duplicates cannot be detected in the response data. Raffle entry records, which cover 47 percent of respondents, contain no duplicate email addresses and no duplicate IP addresses. This is partial rather than complete assurance.
Straight-line responses. Not screened in this cycle.
Minimum reportable cell size. Adoption counts are reported for all tools. Mean ratings are reported only for tools with 10 or more users in a category. Nineteen of 39 named tools fall below this threshold; their adoption counts appear in the results and their ratings are suppressed. No category leader is affected. The threshold is applied mechanically rather than case by case.
Selection. The sample is drawn from one media audience and cannot be generalized to all DPC physicians. Physicians who follow a DPC-focused podcast and newsletter may be systematically more engaged with tooling decisions than those who do not.
Practice size skew. Solo practices are 69 percent of the sample. Multi-physician practice findings rest on 41 respondents across three size bands.
Incentive effects. Response volume was 3.3 times higher during the raffle window than outside it, so respondents recruited under the incentive could in principle differ from those who responded without one. This was tested; see section 8b. No difference was detected in sample composition, tool selection, or ratings, within the limits of what a sample this size can detect.
Response quality under incentive. Raffle entry rewarded completion rather than accuracy, and no straight-line or duplicate screening was performed.
Users rate only their own tools. Ratings reflect satisfaction among adopters. Physicians who evaluated and rejected a tool are not represented in that tool's ratings, which biases all ratings upward relative to the full population of people who considered each product.
Incomplete option lists. Documented in section 6.
No importance weighting. Documented in section 5.
Single wave. No repeated measurement and no ability to observe change over time. The 2025 Battle of the EHRs study used a different instrument built independently, so the two studies are not comparable at the item level.
Reproducibility. The analysis below is reproducible from the raw export using incentive-comparison.py.
The raffle covered 20 of the 63 fielding days, which creates a natural comparison: the same instrument and the same audience, with and without an incentive available. Eighty respondents fall inside the window and 51 outside it.
Sample composition. No detected difference in practice size (p = 0.46), years operating in DPC (p = 0.93), active member count (p = 0.61), or self-reported recruitment source (p = 0.12).
California concentration. The prize was Connection Pass points for a California summit, which raised the possibility that the incentive drove California's position as the largest state in the sample. It did not. California is 18 percent of in-window respondents and 16 percent of those outside it (Fisher exact p = 1.000). The California concentration is a property of the audience, not an artifact of the prize.
Tool selection. No detected difference in the distribution of tools selected in any of the nine categories (all p ≥ 0.06; the lowest was labs and imaging at p = 0.060).
Ratings. Compared at the respondent level, using each respondent's mean across all rating items they answered:
| n | Mean | SD | |
|---|---|---|---|
| In-window | 77 | 4.52 | 0.48 |
| Outside | 49 | 4.42 | 0.83 |
Difference +0.10 on a 1 to 5 scale. Mann-Whitney p = 0.96, Welch t-test p = 0.45, Cohen's d = 0.16. Respondents in the two groups also answered a similar number of rating items (15.5 versus 16.2, p = 0.19).
A false lead, recorded because it is instructive. Comparing the two groups across all 18 category-metric pairs, 17 of 18 differences were positive, which a sign test flags at p = 0.0001. That result is an artifact. The 18 comparisons are not independent, because the same respondents contribute to many of them, so a handful of consistently generous raters produces a run of same-signed differences. The respondent-level comparison above is the correct test, and it shows no difference. The category-level pattern is reported here so that it is not rediscovered later and mistaken for a finding.
What this rules out, and what it does not. With 77 against 49 respondents, this comparison has 80 percent power to detect a difference of about 0.33 rating points. It therefore rules out a moderate or large incentive effect on ratings. It cannot rule out a small one. The observed difference of 0.10 points is well inside the range that would go undetected at this sample size.
The conclusion is that no incentive effect large enough to change a category outcome was detected, not that no effect exists.
This section is a working record. Item 1 is resolved; the remainder will be settled before the 2027 instrument is fielded.
Raw response data is retained by My DPC Story. The 2025 Battle of the EHRs raw export is not available; that study is described in this appendix at the instrument level only, and its published figures are not reproducible from source.
Survey instruments for both studies were reviewed by Maryal Concepcion, MD FAAFP, founder of My DPC Story, for relevance to DPC physicians. She contributed questions and suggestions. Research design, analysis, and written conclusions are the author's.