Downloadable evidence data

PlantLightIndex Data Hub

Download the structured records behind the plant database and fixture measurement layer without losing evidence scope, confidence or source provenance.

Houseplant and grow-light datasets connected to source evidence, PPFD measurements and confidence fields
The public datasets preserve null values and evidence context rather than exporting numbers alone.
Plant dataset

30 plant records

Practical and contextual PPFD, taxonomy, recommendation mode, evidence scope, confidence and source identifiers.

Open plant dataset
Fixture dataset

14 fixtures · 17 PPFD points

Product specifications, data status, confidence, distance eligibility, measurement series and provenance.

Open fixture dataset

PlantLightIndex publishes the data behind its tools

The Data Hub provides machine-readable access to the two datasets that power the current product: houseplant light references and grow-light fixture measurements. The downloads are not scraped lists of generic care numbers. They preserve recommendation mode, evidence scope, confidence, source identifiers and missing values so the uncertainty visible on the website remains visible in the data.

The datasets are intended to make the site auditable and easier to reuse for analysis. A CSV row can be traced back to a canonical plant or fixture page and then to the source registry.

Measured values, estimates and recommendations are different things

PlantLightIndex treats a measurement as an observation, not automatically as a recommendation. A PPFD reading taken at a leaf tells you how much photosynthetically active photon flux reached that position at that moment. A published greenhouse treatment tells you what researchers supplied under a particular experiment. A photosynthetic light-saturation point describes a physiological response. A practical indoor reference is an editorial recommendation intended to help a houseplant grower make a decision. Those records can inform one another, but they are not interchangeable.

This distinction prevents a common failure in plant-light advice. A plant may survive under a very low PPFD treatment while growing slowly or changing form. Another study may show photosynthesis continuing to increase until a much higher PPFD. Neither number, by itself, proves that the lower value is ideal or that the higher value should be used in a home. PlantLightIndex therefore stores measurement context, evidence scope, source type and recommendation status alongside the number.

Missing evidence stays visible

Some plant profiles contain a practical numerical PPFD range. Others show only a parent-species or genus reference. Some remain qualitative because a defensible species- or cultivar-specific numerical target has not been established in the sources reviewed for the site. That unevenness is intentional. A database becomes less useful when every empty field is filled with a plausible-looking number.

When a value is absent, the site should explain what is known instead. A page may state that a plant is commonly described by an extension source as preferring bright indirect light, or that a cultivar has demonstrated a different physiological response from its green parent, without converting those statements into a precise household target. The calculators then change behavior according to the record: full, contextual or qualitative.

Houseplant Light Dataset

The plant dataset contains the 30 published plant and cultivar entities currently used by the database. It includes accepted scientific identity where available, cultivar and entity type, practical PPFD range, contextual range, physiological fields when stored, DLI target status, calculator mode, recommendation basis, evidence scope, confidence and the source identifiers tied to the profile.

Null fields are meaningful. They indicate that PlantLightIndex has not assigned that value under the current evidence rules. Users should not fill those values by copying a range from another species or by converting qualitative wording into a number.

Open the Houseplant Light Dataset →

Grow Light Fixture Dataset

The fixture dataset contains the current fixture catalog plus the measurement status needed to understand whether a product can power distance matching. It includes brand, model, fixture type, power, PPF where published, spectrum fields, data status, confidence, distance-finder eligibility, measurement count, stored PPFD-distance observations and source identifiers.

A spec-only fixture remains in the dataset because product specifications are still useful. Its missing PPFD curve is preserved as missing rather than estimated from wattage or PPF.

Open the Grow Light Fixture Dataset →

CSV is the first public distribution

CSV is widely readable in spreadsheet software, programming languages and data tools, so each canonical dataset page provides a direct CSV download. The file is generated from the same version-controlled records used by the website. This reduces the risk that a manually maintained download drifts away from the live product.

Future versions can add JSON or additional distributions if there is a real use case. PlantLightIndex does not advertise a format that does not actually exist.

Dataset structured data is limited to canonical dataset pages

PlantLightIndex uses Dataset structured data on the dedicated landing pages rather than marking every chart or database view as a separate dataset. Google describes Dataset markup as metadata for a dataset and supports distributions such as CSV through DataDownload. It also recommends provenance properties such as isBasedOn when a dataset aggregates or transforms multiple originals.

See Google Search Central's Dataset documentation and Schema.org Dataset for the underlying model. The site schema points to the real CSV endpoint and lists source URLs as provenance.

Provenance stays attached to the data

Each dataset contains source IDs rather than stripping the evidence relationship away. The landing-page structured data additionally uses source URLs as isBasedOn provenance. A user can therefore move from a CSV row to the Source Library and then to the original publication or manufacturer page.

This is especially important because the datasets combine several kinds of evidence. A plant practical target, a cultivar saturation point and a fixture PPFD measurement can all use the unit µmol/m²/s while representing different claims.

How to use the downloads responsibly

  1. Preserve null values as unknown or not established.
  2. Keep practical and contextual PPFD fields separate.
  3. Use evidence scope and recommendation basis when filtering plants.
  4. Do not treat physiological saturation points as household targets.
  5. Do not use fixture wattage or PPF to infer missing PPFD.
  6. When using measurement rows, preserve distance and dimmer context.
  7. Re-check the source registry when a decision depends on one value.

Versioning and future expansion

The website's source data lives in version-controlled TypeScript records today. As the database grows, PlantLightIndex may migrate storage to a database, but the public concepts should remain stable: identity, observations, evidence claims, recommendation layer, confidence and provenance.

Additional plants or fixtures automatically become useful only after their validation fields are complete. The data validator checks duplicate slugs, broken source links, unsupported full-mode targets and other integrity conditions before production builds.

What the Data Hub is not

It is not a claim that all values are experimentally proven optima. It is not a substitute for measuring a user's room. It is not a warranty that a manufacturer fixture performs identically in every setup. It is a transparent snapshot of the evidence-aware records currently used by PlantLightIndex.

For the rules behind those records, begin with the Methodology Hub. For original references, use the Source Library.

Why the datasets are split into plant and fixture layers

Plants and fixtures answer different sides of the same lighting problem. The plant dataset describes biological or horticultural references and their evidence. The fixture dataset describes devices and measured photon delivery. Combining them into one giant table would repeat data and blur the distinction between a target and a measurement.

The matching tools join the layers only when both sides contain compatible numerical information. This normalized design is easier to audit and expand.

Machine-readable data should preserve editorial meaning

A download can lose important context if it exports only names and numbers. PlantLightIndex includes fields such as calculator mode, recommendation basis, evidence scope, confidence and data status because those attributes determine how the number should be used. An analyst should be able to reproduce the same caution shown on the website.

This is also why missing values are exported as empty fields rather than substituted with zero. Zero PPFD means darkness; an empty practical PPFD means no target is currently established. Those states must remain distinct.

Canonical landing pages explain the distributions

The CSV endpoints are files for machines and spreadsheets. The landing pages are for interpretation. Each landing page explains columns, provenance, limitations and recommended uses, and it is the canonical page carrying Dataset structured data. This avoids creating multiple competing dataset descriptions across charts and directory pages.

The charts can still visualize the same records, but they are treated as interfaces to the data rather than separate canonical datasets.

Future data formats should solve a user need

JSON, APIs or bulk source exports may become useful as the database grows. PlantLightIndex should add them when users or internal tooling need them, not simply to create more technical-looking endpoints. Any new distribution should be generated from the same validated records and documented on the canonical landing page.

Until then, CSV provides a simple, durable format that can be inspected without special software.

Why a data download is useful even for non-developers

A CSV allows growers to sort plants by evidence mode, compare confidence, identify which species have practical PPFD references and keep a private measurement sheet. It also lets users compare fixture measurement coverage without depending on the order of cards on the website. The files open in ordinary spreadsheet software.

The important part is that the download does not strip away the evidence fields. Users can see which rows are strong numerical records and which remain contextual or incomplete.

Why CSV endpoints are generated rather than uploaded manually

A manually edited spreadsheet can drift away from the production code. By generating the distribution from the same arrays used by the site, PlantLightIndex reduces that divergence. A code release that changes a plant range, fixture point or source relationship changes both the web interface and the download together.

This approach also makes QA easier because row counts can be compared directly with published entity counts.

Data-hub audit checklist

  • The canonical landing page exists and is indexable when production is indexable.
  • The CSV endpoint returns a real text/csv response.
  • The Dataset schema points to that real distribution.
  • The dataset description is plain text and within supported limits.
  • Source provenance is represented with isBasedOn URLs.
  • Null values remain null/empty, not zero.
  • Chart pages do not duplicate the canonical Dataset markup.
  • The sitemap includes the landing page, not every internal query variation.

This keeps the data layer both useful to people and consistent with search-engine dataset discovery guidance.

How the data layer supports future products

The public datasets make it possible to build new PlantLightIndex tools without creating a second, inconsistent data source. A future room-light history feature, fixture comparison chart or research-gap dashboard can read the same plant modes, confidence fields and measurement records. Product expansion therefore starts from the evidence model rather than from isolated page copy.

A larger dataset may eventually justify an API. If that happens, API responses should preserve the same concepts as the CSV: null values, evidence scope, recommendation basis, measurement status and provenance. A convenient endpoint that strips away uncertainty would be easier to consume but less faithful to the product.

The Data Hub also creates a clear boundary between editorial content and data. Articles explain concepts and decisions. Canonical dataset pages document the distributions. Calculators apply the records. Keeping those roles distinct reduces duplicate intent and makes the site architecture easier for users and search systems to understand.

How dataset provenance differs from ordinary page citations

An article citation supports a sentence or explanation. Dataset provenance describes where the structured collection came from. PlantLightIndex aggregates taxonomy sources, extension guidance, research observations and manufacturer measurements into its own reviewed records, so the canonical dataset pages identify the underlying source URLs using isBasedOn while the CSV carries stable source IDs.

This distinction matters when the data are reused outside the website. A person can import the CSV without copying all source URLs into each row, then look up only the records relevant to the analysis. Search systems can read the canonical dataset metadata and see that the distribution is derived from multiple originals rather than being presented as raw experimental data collected entirely by PlantLightIndex.

The landing pages also explain the transformation that occurs between source and dataset. PlantLightIndex normalizes plant identity, separates practical and contextual ranges, labels evidence scope and converts fixture observations into a consistent measurement structure. Those editorial transformations are why provenance and methodology must accompany the download.

As the database grows, the Data Hub should remain the canonical entry point for machine-readable material. New visualization pages can link back here rather than each claiming to be a separate dataset. This reduces schema duplication and gives users one stable place to understand field meanings and downloads.

Dataset landing pages should also state what the files do not contain. The current downloads are evidence-aware reference datasets, not complete plant-care profiles, raw laboratory data or standardized independent fixture tests. Naming those boundaries helps prevent secondary users from applying the CSV to questions it was not designed to answer.

Users who create derived datasets should keep a clear distinction between original PlantLightIndex fields and their own transformations. For example, a modeled PPFD estimate, custom ranking score or inferred DLI target belongs in a new derived column rather than replacing the published field. That simple practice makes later updates easier to reconcile and keeps the provenance chain intact.