# RxID Commons: research brief

_Prepared for RxAI · August 28, 2026_

## Executive conclusion

There is a credible opening for an open drug **image and imprint identity layer**, but not for a day-one open replacement of the full Medi-Span/FDB/Gold Standard stack.

The incumbents do much more than aggregate manufacturer files. They normalize identifiers, maintain proprietary product hierarchies, edit pricing benchmarks, author clinical content, and operate drug-utilization-review (DUR) rules. Their public claims also complicate a blanket “slow to update” story: Elsevier advertises daily Gold Standard updates; Medi-Span offers daily or faster pricing updates in some tiers; FDB exposes weekly new-drug additions between monthly MedKnowledge updates. The sharper problem is that the best image/imprint data is closed, redistribution is restricted, provenance is hard for outsiders to inspect, and public image infrastructure has regressed since NLM retired Pillbox.

The most defensible wedge is therefore:

> An open, versioned, provenance-rich registry of solid-dose appearance—linked to NDC and RxCUI, seeded from federal data, updated directly by labelers, and continuously validated against consented real-world pharmacy images.

That product can become the reference layer for pill identification models, pharmacy verification tools, poison-control workflows, accessibility products, and developer applications. Clinical screening, reimbursement, and pricing editorial products should be treated as later and separate markets.

## 1. Incumbent landscape

### What the vendors sell

| Vendor | Core products and relevant capabilities | Update / delivery claims | Takeaway for RxAI |
| --- | --- | --- | --- |
| **Medi-Span (Wolters Kluwer)** | MED-File product hierarchy and identifiers; Drug Image Database v2.0; Drug Imprint Database v2.0; pricing (AWP, WAC, DP, FUL, NADAC and other benchmarks); clinical screening/DUR; patient education; flat files, APIs, and web services. Medi-Span publishes “more than 20,800 unique images” mapped to more than 56,600 NDCs and imprint records for more than 73,800 NDCs. | Pricing products advertise daily or multiple-times-per-day updates depending on tier. | Broad coverage and mature integration. The opportunity is openness, traceable source evidence, field-image variation, and machine-learning rights—not claiming Medi-Span has no images. |
| **First Databank (FDB)** | FDB MedKnowledge descriptive/product data, images and imprints, clinical modules, interaction/contraindication screening, patient education, terminology/interoperability modules, and drug price types including WAC, DP, SWP, FUL and NADAC. | FDB MedKnowledge Explorer says newly added drugs are available through weekly automatic updates between monthly updates. | A strong incumbent for embedded drug knowledge. Manufacturer-reported pricing and SPL-derived materials demonstrate that direct-source contribution is already part of the market’s supply chain. |
| **Elsevier Gold Standard Drug Database** | Product data and images, pricing, reference content, clinical decision support/DUR, prescriber order entry, and patient education. Elsevier advertises a modern API/relational delivery model and 32 hierarchy levels. | Elsevier states that Gold Standard data is delivered daily. | Again, “always current” alone is not enough. RxAI must differentiate through open licensing, transparent versions, contribution workflows, and image validation. |
| **Oracle Health Multum (formerly Cerner Multum)** | Multum MediSource Lexicon / drug database: normalized generic drug identifiers, NDC product descriptions, ingredients, dose forms, strengths and relationships. Multum is also a source vocabulary in RxNorm and underlies consumer drug information products. | Current public commercial detail is thinner than for the other vendors; NLM documents the files supplied to RxNorm. | Useful incumbent/terminology comparator, but not the strongest image-specific benchmark. Avoid implying that RED BOOK/Micromedex and Multum are the same product: they are distinct sources even when surfaced together by downstream sites. |

Primary product sources: [Medi-Span content sets](https://www.wolterskluwer.com/en/solutions/medi-span/medi-span/content-sets), [Medi-Span overview](https://www.wolterskluwer.com/en/solutions/medi-span), [FDB MedKnowledge Explorer](https://www.fdbhealth.com/solutions/fdb-medknowledge-explorer), [FDB pricing policy](https://www.fdbhealth.com/drug-pricing-policy), [Elsevier Gold Standard](https://www.elsevier.com/products/gold-standard-drug-database), [NLM Gold Standard source synopsis](https://www.nlm.nih.gov/research/umls/rxnorm/sourcereleasedocs/gs.html), and [NLM Oracle Health Multum source documentation](https://www.nlm.nih.gov/research/umls/rxnorm/sourcereleasedocs/mmsl.html).

### Pricing and procurement evidence

The vendors generally use quote-based pricing, so “a pharmacy pays $X” is not a reliable universal statement. Public procurement records do, however, substantiate five- and six-figure institutional spend:

- A 2025 Florida procurement notice describes 15 Medi-Span Price Rx licenses totaling **$180,357.81 over three years** (about $60,119 per year across the license group). This is an institutional pricing reference, not a retail-pharmacy list price. [HigherGov procurement summary](https://www.highergov.com/sl/contract-opportunity/fl-medi-span-price-rx-59416232/)
- A 2026 Utah renewal quote shows the **FDB MedKnowledge Enhanced Package base fee at $17,602 in year one**, before additional packages and user fees. [GovTribe document summary](https://govtribe.com/file/government-file/4-fdb-quote-redacted-dot-pdf)
- A UK public contract awarded **£185,367 over three years** for FDB licenses. [UK Contracts Finder](https://www.contractsfinder.service.gov.uk/notice/0e9f6d77-b0fa-4d7d-bc78-5c1c5b438598)
- U.S. federal notices repeatedly describe FDB purchases as sole-source or brand-name requirements, evidence of both integration depth and switching cost. [SAM.gov MedKnowledge notice](https://sam.gov/opp/2f81729f77e04b43a0f0d16fa22984a3/view)

The responsible marketing claim is: **commercial drug-data access can cost institutions tens of thousands of dollars per year or more, depending on scope and users**. It should not be phrased as a universal price paid by every pharmacy.

### Pain points: what is supported and what needs validation

**Supported by public evidence**

- **Closed redistribution rights.** Medi-Span’s public terms reserve proprietary rights and restrict reproduction, transmission, collective works, and resale absent permission. [Medi-Span Price Rx terms](https://www.wolterskluwer.com/en/solutions/medi-span/about/price-rx-terms-use-disclaimer)
- **Opaque provenance and transformation.** Vendors describe editorial processes, but outsiders generally cannot inspect a public chain from manufacturer submission → edit → canonical image → downstream release.
- **High switching cost.** Proprietary IDs, product hierarchies, clinical logic, and embedded integrations make replacement hard even when public identifiers exist.
- **Public image infrastructure was discontinued.** NLM retired Pillbox, its image library production, APIs, and search experience in January 2021. The final files remain available but are explicitly stale and “should not be used for pill identification.” [NLM retirement notice](https://www.nlm.nih.gov/pubs/techbull/ja20/ja20_pillbox_discontinue.html)

**Claims that should be tested with customer interviews**

- Incumbent image gaps by dosage form, labeler, or launch cohort.
- Time from a manufacturer’s change to availability in a specific customer feed.
- How often pharmacy systems receive and load vendor updates, which is different from how often a vendor publishes them.
- Frequency and cost of incorrect imprint, shape, color, or NDC mappings.
- Whether buyers will adopt an open image/imprint layer alongside—not instead of—their clinical knowledge base.

## 2. Public/free seed sources

### DailyMed and FDA Structured Product Labeling (SPL)

DailyMed is the best starting point for labeler-submitted product facts. SPL can encode NDCs, ingredients, strengths, dosage forms, shape, color, score, size, imprint, and an optional solid-dose image. DailyMed offers daily, weekly, monthly, and full label downloads plus web services. It does **not** review SPL content before publication, and “in use” labeling may differ from FDA-approved labeling; provenance and verification flags are therefore essential. [DailyMed about](https://dailymed.nlm.nih.gov/dailymed/about-dailymed.cfm), [download resources](https://dailymed.nlm.nih.gov/dailymed/spl-resources.cfm), [SPL image guidelines](https://dailymed.nlm.nih.gov/dailymed/splimage-guidelines.cfm)

DailyMed’s image specification is useful as an acquisition baseline: obverse and reverse views, a 1024×768 JPEG, controlled gray background, scale references, accurate color/white balance, and preserved dimensions. Manufacturer submissions are the authoritative declared reference, but they are not sufficient evidence of how a product appears across production lots, cameras, lighting, wear, or dispensing environments.

### RxNorm and RxNav APIs

RxNorm supplies normalized drug names and RxCUIs and connects names from multiple vocabularies. It is ideal for the semantic spine above NDC/package-level appearance records. NLM releases the full set monthly and weekly SPL-derived additions; RxNav/RxNorm REST APIs offer programmatic lookup. RxNorm normalized names/codes are public domain, but the **full** distribution contains proprietary source vocabularies with separate restrictions. An open project should ingest the public RxNorm/SPL subset or carefully filter by source, not republish all full-release atoms. [RxNorm overview](https://www.nlm.nih.gov/research/umls/rxnorm/overview.html), [terms](https://www.nlm.nih.gov/research/umls/rxnorm/docs/termsofservice.html)

### FDA NDC Directory and openFDA

The FDA NDC Directory is the daily-updated product/package registry for electronically listed final marketed drugs. It supplies labeler, product, package, ingredient, strength, dosage-form, route, marketing, and application fields. FDA warns that listing does not mean approval or verification and that the labeler is responsible for submitted content. openFDA exposes this data in JSON and bulk downloads, with harmonized identifiers where possible. openFDA content is generally public domain/CC0, but its terms note that some third-party works can retain rights. [openFDA NDC API](https://open.fda.gov/apis/drug/ndc/), [NDC data description](https://open.fda.gov/data/ndc/), [openFDA terms](https://open.fda.gov/terms/)

### NLM Pillbox: the precedent and the gap

Pillbox combined SPL-derived physical characteristics, NDCs, RxNorm mappings, and NLM images into a public pill-identification resource. NLM retired the search sites, production dataset, image library, and APIs in January 2021 as part of portfolio consolidation. The final static files remain available for research/development/education, but NLM says they are not updated and should not be used for pill identification. DailyMed subsequently removed RxImage-provided pill images; only images submitted with SPL remain there. [NLM retirement notice](https://www.nlm.nih.gov/pubs/techbull/ja20/ja20_pillbox_discontinue.html), [Data.gov catalog entry](https://catalog-old.data.gov/dataset/pillbox-retired-2f7f9), [RxNav release note](https://lhncbc.nlm.nih.gov/RxNav/?id=rximageapi_support)

The final Pillbox metadata snapshot is also a useful historical coverage benchmark. A query run for this brief against NLM’s archived Socrata dataset returned **83,925 appearance records**, of which **9,779 were flagged `has_image=True`** (11.7%). It also returned 81,978 records with an imprint value and 19,171 distinct imprint strings. These are 2020-era archival counts, not present-day coverage claims. Dataset: [NLM Pillbox archived data](https://datadiscovery.nlm.nih.gov/d/crzr-uvwg).

Pillbox proves both demand and fragility: a public project can acquire high-quality standardized images and enable research, but a single-agency service without a durable contribution, funding, and governance model can disappear. RxAI should design for exportability, mirrors, versioned releases, and governance independence from the start.

## 3. FDA imprint requirements and why imprints are not standardized

21 CFR Part 206 requires most solid oral dosage forms to bear a code imprint. Under §206.10, the code **together with size, shape, and color** must permit unique identification of the product, manufacturer/distributor, active ingredient(s), and strength. A code can be a letter, number, word, company name, NDC, symbol, logo, monogram, or combination assigned by a drug firm. Letters/numbers are encouraged but not mandatory. Exemptions exist, including some investigational, compounded, radiopharmaceutical, technologically infeasible, and controlled-setting products. [21 CFR §206.10](https://www.law.cornell.edu/cfr/text/21/206.10), [CFR Part 206 text](https://www.govinfo.gov/content/pkg/CFR-2023-title21-vol4/pdf/CFR-2023-title21-vol4-chapI-subchapC.pdf)

FDA guidance further notes that look-alike products and similar, missing, or illegible imprints have contributed to wrong-product or wrong-strength errors; it recommends differentiation and warns that symbol-only imprints are harder for alphanumeric databases to index. [FDA product-design safety guidance](https://www.fda.gov/media/84903/download?attachment=)

The consequence is important: the regulation specifies a **uniqueness outcome**, not a global imprint syntax or assignment authority. Drug firms choose their own code system, and uniqueness depends on a tuple of visual characteristics—not imprint text alone. Logos can be difficult to serialize; scores can be confused with separators; side order and orientation matter; and the same text can be ambiguous without color/shape/size/manufacturer context. The open schema should model the observation, not flatten it into one string.

Recommended imprint representation:

- separate obverse/reverse faces and reading order;
- literal text, normalized text, and OCR alternatives;
- score count/type and whether a separator is a score or an imprint delimiter;
- vector/symbol region plus human-readable symbol taxonomy;
- emboss/deboss/ink method, orientation, legibility, and confidence;
- size, thickness, shape, color distribution, coating, and capsule-part colors;
- source record, submission date, effective date, lot (when permitted), and supersession history.

## 4. The credible wedge

### Positioning

Do not launch as “free Medi-Span.” Launch as **the open visual identity layer for medicines**.

The wedge works because:

1. Federal sources already provide a daily product spine and manufacturer-declared appearance data.
2. Manufacturers have a safety and support incentive to correct appearance records quickly.
3. Pharmacies possess the real-world image distribution that static reference photos cannot capture.
4. Computer vision needs training/validation rights, provenance, negative examples, and variation—not just a single polished image.
5. Developers can adopt an image/imprint API alongside their existing terminology/DUR vendor, reducing switching risk.

### What manufacturers should contribute

Minimum contribution contract:

- labeler identity and authorization to publish;
- NDC product/package mappings and RxCUI candidates;
- explicit imprint logic by face, including symbol assets;
- shape, dimensions, score, coating, and controlled color reference;
- calibrated obverse/reverse images plus optional multi-angle captures;
- marketing start/end dates, change effective date, and superseded appearance;
- manufacturing/site or product-family context where publishable;
- change webhook/feed and named data steward;
- license covering public redistribution and ML training/evaluation;
- signed attestation plus correction/recall workflow.

High-value later additions:

- approved lot-level variation envelopes and known tooling variants;
- packaging-to-dose lineage and repackager/source-NDC relationships;
- anti-counterfeit or verification signals that can safely be public;
- negative examples and common OCR confusions;
- recall, discontinuation, and temporary-supply-change notifications.

### Real-world validation model

The proposed “center of the distribution becomes canonical” idea is directionally useful but should not reduce truth to a single centroid. Valid products can be multimodal because of manufacturing site, tooling, coating, camera, packaging, aging, or an authorized appearance change.

A safer model:

1. Preserve the manufacturer reference as a declared source, not automatically as ground truth.
2. Accept pharmacy images only with agreements that address PHI removal, retention, chain of custody, consent, and ML rights.
3. Segment the dosage form and normalize scale/color while retaining the raw capture.
4. Cluster by product plus time/lot/site signals where available.
5. Promote **validated appearance modes** after minimum sample, source-diversity, and reviewer thresholds.
6. Treat outliers as “review required,” never as counterfeit or wrong drug without independent evidence.
7. Publish confidence, evidence count, modality, freshness, and version history through the API.

### Synthetic images

Synthetic references can improve coverage across lighting, angle, background, occlusion, wear, and camera conditions, but must never be visually or programmatically confused with observed evidence. Every asset should carry `origin: manufacturer | pharmacy | public_source | synthetic`, a generator/version, parent image IDs, transformations, license, and intended-use restrictions. Synthetic data should augment—not vote in—the validation distribution.

## 5. Recommended phase-one scope

**In scope**

- U.S. human solid oral dosage forms.
- NDC product/package + RxCUI linkage.
- Imprint/shape/color/size search.
- Manufacturer submission portal/feed spec.
- Versioned image store with real/synthetic/source labels.
- Field-image ingestion pilot with 3–5 pharmacy partners.
- Open API, bulk snapshots, and correction workflow.

**Out of scope at launch**

- Consumer-facing definitive pill identification.
- DUR, dosing, interaction, allergy, or contraindication advice.
- AWP/WAC editorial pricing replacement.
- Counterfeit determinations from an image alone.
- Non-U.S. regulatory normalization or non-solid dosage forms.

## 6. Governance, licensing, and safety requirements

- Pick a database license only after counsel reviews source compatibility. A permissive code license does not automatically solve database rights, third-party image rights, trademarks, or labeler-submitted copyrights.
- Separate code, metadata, original images, field images, and synthetic assets into explicit license classes.
- Publish source-level provenance and prevent proprietary RxNorm source atoms from leaking into open releases.
- Use signed releases, checksums, immutable snapshots, mirrors, and an export path so the project survives an operator shutdown.
- Establish a clinical/data governance board and a public correction/SLA policy.
- State clearly that the API supports identification workflows and is not, by itself, medical advice or a sole basis for dispensing.
- Build evaluation sets with hard negatives and report performance by manufacturer, color, shape, imprint type, camera, and subgroup rather than one aggregate accuracy number.

## 7. Name directions

These are creative directions only; trademark and domain clearance are still required.

1. **RxID Commons** — strongest balance of pharmacy meaning, public infrastructure, and extensibility.
2. **OpenImprint** — literal, memorable, and tightly focused.
3. **Doseprint** — suggests a visual fingerprint; more brandable than descriptive.
4. **Tablet Commons** — warm and open-source, but narrower than capsules/other solid forms.
5. **OpenDose Registry** — institutional and credible; less distinctly image-first.
6. **Pillprint** — approachable and visual, though more consumer-sounding.
7. **Pharmark** — compact “pharma + mark”; requires careful trademark screening.
8. **RxSpecimen** — scientific and image-forward, with a slightly colder tone.

The concept site uses **RxID Commons** as a working name.

## 8. Immediate validation questions

Before committing to a full build, interview at least five manufacturers, five pharmacy/dispensing software teams, and five potential API users. Test:

- Which product-change events routinely arrive late or ambiguously today?
- Who inside a manufacturer owns imprint/image data and can sign a public license?
- Can pharmacies legally and operationally contribute de-identified product images at dispense time?
- Which API response fields would let a buyer add this alongside FDB/Medi-Span rather than replace them?
- What evidence and liability terms are required before an image can influence dispensing?
- Will an open bulk snapshot matter more than a free hosted API?
- Which sustainability model—sponsor membership, hosted API tiers, certification, or enterprise support—best protects the open core?

## Bottom line

RxAI should compete first on a layer the incumbents and federal sources do not currently provide together: **open access + direct manufacturer change feeds + observed real-world variation + ML-ready rights + public provenance**. That is a narrow enough promise to ship credibly and a valuable enough data network to expand from later.
