About‍ ➜ Methodology

How Interior Medicine Evaluates Products

Dr. Meg Christensen is the physician founder of Interior Medicine, a non-toxic home resource built on her background in medicine, biochemistry, epidemiology, and clinical research.

Published May 28, 2026 | Updated June 9, 2026

Methodology Overview

Interior Medicine evaluates products to find what is true, not to offer an opinion, perpetuate wellness claims, or restate marketing as fact.

Every product is rated against consistent scales, applied the same way each time. Brand claims are never taken at face value. Products are assessed against the exact same standards: objective third-party certifications, published data, and documented evidence. Held to one consistent standard, the differences between them become clear.

There are two kinds of scales: Material Health scales rate what a product is made of. The Health Function scale rates how well a product performs the specific health job it exists to do, and applies to certain categories. Each scale below links to its full definition:

Material Health Scales

Each is a full guide: the rating scale plus a deep reference on that material's certifications, hazards, and what to look for.

Health Function Scale

One scale, consistent across categories, rating how well a product does the health job it's built for. What earns each tier is defined per category.

How Material-Based Products Are Evaluated

Furniture, Bedding, Cribs, Mattresses, Textiles, Shower Curtains, and Cookware Categories

These products are evaluated based on what hazards they could introduce into close, sustained contact with the body. A couch, a mattress, a crib, a set of sheets, a shower curtain, a pan — these are products spent significant time in contact with, either through skin, through the air in the room they occupy, or through what is cooked or served on them. The relevant question is not whether the finished product is "non-toxic" as a whole (no single certification evaluates a finished couch), but what each material in the product contributes to the surrounding air, dust, and skin contact over time.

Most materials in this category have third-party certifications that test for substances in the hazard database, which makes layer-by-layer evaluation possible. Each material (wood, foam, fabric, glue, metal, ceramic, glass) is evaluated separately against the certifications relevant to that material. Full certification logic and tier criteria for each material are in the Material Health Guides.

Every rating is relative. Nothing is perfectly non-toxic, and nothing is 100% toxic. Each material sits somewhere on a spectrum from healthiest to harmful, defined by the certifications and evidence available at this moment. The spectrum itself shifts over time: materials standard a decade ago have been reformulated or phased out, new certifications appear, existing ones get tightened or watered down, and the research keeps moving. A rating on this site is a snapshot, not a permanent verdict, and the goal of the methodology is to make sure each snapshot is taken honestly and consistently against the best available information.

Every material is then sorted onto the same five-tier scale:

The rating reflects what is known about the material. A Healthiest rating means the strongest available certifications confirm what is excluded. A Harmful rating means the absence of those protections is itself the finding. Each product is taken apart layer by layer, and each layer rated separately.

The Six-Step Framework

This framework comes from toxicology. It's how risk is formally assessed, and it's the structure I apply to every material-based product review on Interior Medicine. An overview of the structure is here.

Most non-toxic resources stop at step 1 — they list hazardous substances and tell you to avoid products that contain them. That approach mistakes hazard for risk, and the two aren't the same. A substance has to be present, reach you, and reach you at a meaningful dose, in a body that responds to it, before it becomes a risk. Skipping the middle steps is what produces both the fearmongering ("everything causes cancer") and the dismissal ("everything is fine") that dominate this conversation. This framework is the alternative.

  1. Hazard. Is the substance harmful at any dose? Answered using a database of 70 substances drawn from regulatory and authoritative bodies (IARC, NTP, REACH SVHC, CLP CMR, ATSDR, WHO, EPA, Prop 65, AB 1200), emerging-concern sources (Green Science Policy Institute's Six Classes, DTSC Candidate Chemicals List, EPA CCL 6), and substances frequently flagged on wellness blogs that institutional lists miss for structural reasons (bleach, ammonia, quats, fragrance) but for which some evidence exists. Hazard is the starting point of the framework, not the endpoint. Full hazard database and references here.

  2. Exposure. Can a hazard actually reach a person from the product? Answered by combining what certifications confirm is excluded or limited (most certifications test extraction or emission, not just content) with chemical behavior: VOCs become gas, sVOCs migrate slowly to dust and surfaces, particles shed with friction, heavy metals stay locked unless heat, acid, or friction acts on them. It is vital to address whether a substance can leave a product at all. Full exposure reasoning here.

  3. Dose. How much exposure actually reaches the body and has the potential to raise the burden in bloodstream or organs? Answered by combining exposure level with mitigation behaviors (ventilation, wet dusting, HEPA vacuuming, keeping hot or acidic food off questionable materials). Dose is where individual behavior changes the answer, which is why two households can own the same product and end up with different exposures. The five mitigation strategies are described here.

  4. Dose-Response. How does the body respond to that dose? Often this isn’t known with precision. The toxicology literature establishes dose-response curves for individual substances at occupational or experimental doses, but residential exposures usually fall well below those levels and almost always involve mixtures rather than single substances. Interior Medicine treats this step with appropriate uncertainty rather than inventing precision that does not exist. Where good evidence exists, it is used. Where it does not, that is stated.

  5. Susceptibility. How does a particular household's vulnerability or resilience change the answer? Answered by the household, not by Interior Medicine. Vulnerability factors include children under six, adults over 80, pregnancy or breastfeeding, asthma or MCS or autoimmune conditions, and chronic kidney or liver disease. Resilience factors include exercise, sleep, social connection, and diet. Full susceptibility reasoning here.

  6. Risk. The synthesis of the first five steps, weighted by personal risk tolerance.

Hazard, exposure, and dose are knowable or partially knowable from outside data, and Interior Medicine evaluates products against those three steps. Susceptibility is partly knowable from outside data and partly individual to a household. Risk tolerance and precautionary stance are individual. Product recommendations on this site are calibrated to steps 1 through 4. Steps 5 and 6 are not stated for any household other than the author's, because doing so would be a misrepresentation.

How Research Quality Is Evaluated

The hazard database, exposure logic, and dose estimates that feed the framework depend on research. Not all research is equally informative, and a single dramatic study should not be held up as proof of harm, nor should a single inconclusive study be held up as grounds for dismissal. Interior Medicine evaluates research based on what model was studied, at what dose, under what conditions, what the authors actually concluded, and what the broader body of evidence says.

Four questions are applied to any single study:

  • What was studied. Humans, animals, or cells in a dish. Cell and animal studies are useful for screening but do not always translate to humans. Observational studies in humans are stronger but cannot prove causation alone.

  • At what dose. Doses 1,000 times higher than typical residential exposure are common in screening studies. An effect at high doses does not automatically scale down to an effect at low doses. Many substances have threshold effects or non-linear dose-response curves.

  • Under what conditions. Sealed chamber at elevated temperature is not equivalent to regular use at room temperature. An effect under extreme conditions is not the same finding under normal ones.

  • What the authors concluded. Paper titles often imply more certainty than the conclusion section supports.

Beyond individual studies, weight is given to the broader body of evidence.

  • Strong evidence. Multiple independent research groups, study types, and regulatory agencies have reached consistent conclusions across countries and years. This is where meaningful health guidance is established.

  • Substantial but contested evidence. A real body of research exists and the hazard is real, but mechanism, threshold, or significance at realistic exposure levels is still debated.

  • Limited evidence. Some research exists but is preliminary, narrow in scope, or conducted in models distant from human biology.

  • Data gap. No meaningful research exists on the specific question. Absence of evidence is treated as a data gap, not as reassurance or cause for dismissal.

The framework treats each step (hazard, exposure, dose, dose-response, susceptibility) as having its own evidence depth, because the research base is rarely uniform across all five for any given substance. PFAS has decades of consistent research across all steps. A newer phthalate alternative might have strong hazard evidence but limited exposure data and almost no dose-response work at residential levels. Both situations get reported honestly rather than collapsed into a single rating.

When the evidence is strong, the methodology states that. When it is weak or absent, the methodology states that too. Acting on uncertain evidence is a legitimate precautionary choice, but only when the uncertainty itself is named rather than hidden.

The full version of this approach is here.

Health Function Scales

Most products are rated on their material health: what they are made of, and how those materials affect the people living with them. This is the core of the Method and applies to nearly everything in a home.

Some products are also rated on a second scale: their health function. A product can be made of clean materials and still fail at the health job it exists to do, and that failure carries its own cost. A blackout curtain that only dims a room doesn't protect sleep the way full darkness does. A shower curtain that holds moisture invites mold. A radon detector that isn't independently tested can miss the thing it's meant to catch. The Health Function scale rates how well a product performs that job, on a five-tier scale from Highest to Lowest that stays consistent across every category.

When a product is rated on both scales, the two ratings are independent and can point in different directions. An organic cotton blackout curtain may earn the highest material rating while landing lower on health function, because natural fibers can't achieve total blackout on their own. A synthetic curtain may be the reverse. Showing both, rather than collapsing them into a single verdict, is what lets you weigh the tradeoff against your own priorities.

See the full Health Function scale for how each tier is defined.

Two Rating Systems

The framework above is universal, but the specific criteria differ by category, because the relevant question changes with the product. Material-based products are judged on what each layer contributes to air, dust, and skin contact. Water filters and air purifiers are judged on certified performance rather than housing material, since the contaminants they remove are the exposure being addressed. Radon monitors, humidifiers, candles, lighting, and the rest each have their own standard, documented on their own page.

What stays the same across all of them is the logic underneath: every product in a category is judged against the same objective standards, applied equally, so that what surfaces is the truth about the product rather than the strength of its marketing. The specific criteria for each category, including any Health Function ratings that apply, are documented in the FAQ on that category's page.

A Note on “Non-Toxic” Terminology

The words non-toxic, chemical-free, toxin, and toxic appear throughout Interior Medicine despite being scientifically imprecise. No agreed-upon definition of non-toxic exists. Everything is made of chemicals, including water, so nothing is truly chemical-free. Toxin refers specifically to a natural poison or venom, while toxicant is the accurate term for synthetic chemicals with negative health effects. A substance being toxic at some dose does not mean it presents a health risk at the doses encountered in homes.

These terms are used anyway, for practical reasons. They are the most culturally recognized and searchable language for what people are looking for when they research healthier home products. Using technically correct terminology that no one searches for would make the information unfindable. Non-toxic is treated here as shorthand for a complicated problem, not as a literal claim. The terminology will be updated when the cultural language shifts.

Limits of This Methodology

The methodology below reflects the most rigorous approach currently possible given available evidence, certifications, and solo-operator resources. The limits noted here are documented so they can be addressed as the field develops, as new certifications become available, and as the site's capacity expands. Updates will be made as warranted.

  • The hazard database is curated, not exhaustive. It draws from authoritative sources but reflects a working set of 70 substances rather than the full universe of substances potentially present in home products. Substances not yet classified by a source on the list, substances flagged primarily in non-English literature, and substances in early-stage research may be relevant to a given product but absent from the database.

  • Certification reliance is partial coverage, not full disclosure. Where a certification tests for a defined set of substances, the evaluation is robust within that set. Substances outside the certification's scope remain unknowns, regardless of whether the product is otherwise well-certified. A GOTS-certified cotton fabric has been evaluated against GOTS criteria, not against every substance that could theoretically be present. This is the best available proxy for material composition disclosure, but it is a proxy.

  • Material-layer evaluation does not capture interactions between layers. The methodology rates each material layer separately and assigns the product's overall rating to its weakest layer. This is a reasonable simplification but does not account for chemical interactions across layers, the way certain finishes alter the off-gassing behavior of underlying materials, or cumulative effects from multiple low-rated layers.

  • The rating tier system is ordinal, not quantitative. Healthiest, Healthy, OK, Use Caution, and Harmful are relative positions on a scale, not measured exposure values. Two products in the same tier are not necessarily equivalent in actual exposure delivery, and the boundaries between tiers reflect judgment about which certifications and material characteristics warrant which classification rather than measured thresholds.

  • The methodology is documented but not externally audited. No third party has reviewed the rating logic for consistency, the database for completeness, or the per-category criteria for methodological soundness. This is standard for solo-operated resources but is a limit relative to formally peer-reviewed or institutionally audited methodologies.

Where to Go Next

About‍ ➜ Methodology