Longitudinal patient records from German ambulatory care, structured and made accessible for efficient research.
Every answer traces back to one document, on one page, written by the physician who treated the patient.
The documented reason a therapy was started, changed or held.
Confirmed findings and excluded findings, both recorded, both searchable.
General practice records going back years before a specialist opens the file.
One patient, as the practice documented them. Every value carries the document it was read from, the page, and whether it was reviewed.
Green is what the physician confirmed. Red is what the physician excluded. Both are in the file, and both are searchable.
A hospital dataset holds the episode. A claims database holds the invoice. The ambulatory file holds the years in between.
This filter runs on the primary documentation, including the free text in it.
An answer you can take into your next internal meeting, with the reasoning attached.
Cohort sizes against your criteria, distributions, and completeness per variable, so you can see whether an endpoint is supportable.
How we arrived at each number, written in the same document as the number.
So you can present it without rebuilding it.
Know your cohort before you start. We run your criteria against documented records, criterion by criterion, and show the drop-off at each step. Enrolment planned against documentation, not estimates.
When an endpoint depends on a variable your current source does not carry, we tell you on the first call whether the physician documented it, and where.
Claims data shows you that something happened. The clinical record shows you why, in the documentation itself, with the source behind every value.
Tell us your question. On the first call we tell you whether the data can answer it. The counts follow once we have your criteria.
A named person who can read your question answers it.
We tell you what the data supports for your specific question.
Under NDA you see real structured records for your cohort, with a link to the source for every value.
Records are de-identified at source, inside the practice. No personal data ever reaches our infrastructure.
It happens in the practice, under the physician's duty of confidentiality, before any record enters our infrastructure.
De-identification happens before any transmission. We never receive a key, and we never hold personal data.
Art. 9(2)(j) GDPR, § 27 BDSG, § 6 GDNG, under Art. 28 GDPR, with an external legal opinion. Data residency Germany.
Every study gets a protocol and a statistical analysis plan before any data is touched.
Most recruitment plans rest on two inputs that were never checked against data: a site feasibility questionnaire answered from memory, and a prevalence estimate applied to a catchment population. Both tend to be optimistic in the same direction, and the gap shows up once screening has started and the timeline is already committed.
There is a third reason that is harder to see. Everything needed to apply the criteria sits in the file: the documented reason a therapy changed, whether a condition was actively ruled out, a laboratory trend rather than a single value. Almost none of it is coded, so a site has to read it out of the records by hand, which is accurate and slow. The pool that gets reviewed stays small, and a patient who is only temporarily out of range, and who would qualify after one more laboratory test, is never seen.
In Germany this matters more than elsewhere, because chronic disease is largely managed in ambulatory practice and the years before a specialist diagnosis are documented there. Checking the protocol against records rather than against a questionnaire gives you counts per criterion, the drop-off at each exclusion, and the assumptions written down.
This is where MPIRIQ is at its strongest. The cohort usually exists and the sites are usually enrolling. What is needed is a variable: the documented reason a therapy changed, whether a condition was actively excluded, a laboratory trend rather than a single value, the years of history before the referral. All four are written down in the primary documentation, which is the layer we read.
Of the routes available at that point, reading the primary documentation is the one that answers questions of this kind. Chart review at the sites is accurate but consumes exactly the site capacity that is already your constraint. Licensing a second claims dataset adds patients. It does not add the variable you are missing.
The first conversation establishes what is documented for your specific question, so you plan against a known answer rather than an assumption.
Four types, and they answer different questions.
Statutory health insurance claims. The broadest coverage in Germany and the standard source for incidence, prevalence and treatment volumes. Built to settle invoices, so it records that something was billed. Diagnoses carry billing intent.
Hospital and university datasets. Deep for the episode of inpatient care and strong where a condition is managed in hospital.
Disease registries and cohort studies. Purpose-built, well defined variables, high data quality per patient, covering the patients who were enrolled.
Ambulatory practice records. The file the physician keeps in order to treat. Complete, because it was written for care rather than for research. This includes general practice and specialist practice, which is why we say ambulatory care rather than primary care.
The fourth is the layer MPIRIQ works in, structured and coded at source with every value linked back to its document.
Your team checks the values, not our claim about them. Each structured field keeps a link to the document it was read from, the page it appears on, and whether it was reviewed. The record was written and reviewed by the treating physician, and that link back to it is what you check.
Two properties matter most for feasibility work. Completeness per variable, reported per cohort, because a variable that is 40 per cent complete will not carry an endpoint however accurately it was extracted. And whether exclusions are recorded, because a criterion that was actively ruled out is a different thing from one that was never coded, and the primary documentation distinguishes the two.
For formal sensitivity and specificity figures, tell us at the first call which standard your regulatory colleagues work to and we will establish what your submission needs.
Timeline depends on four things: how many patients, how deep the extraction goes, over what observation period, and which variables your endpoint requires. We scope all four in the first conversation.
What we commit to is the time to certainty. A named person replies to your question. A call establishes what the data supports. Under NDA you then see real structured records for your own cohort and check them yourself. None of that depends on scope, so none of it is an estimate.
Study delivery is scoped in writing against what the initial review showed, so the timeline you agree to is the timeline you get.
Send us one indication and one question. Under NDA you see real structured records for your own cohort, with a link to the source for every value, before anyone commits budget.