Data dictionary
Field definitions for the public research edition behind the observatory. Compact CSV and JSON exports are public. Complete evidence ledgers, monitoring, exports, and interpretation are available through the professional pilot. Each release carries an asOf date and is rebuilt weekly.
research-edition.csv columns
| Column | Type | Description |
|---|---|---|
company | text | Issuer name as it appears in the source filing. |
exchange | text | TSX, TSXV, or Unmatched (matched to the listed universe by name). |
symbol | text | Ticker on the matched exchange, where resolved. |
year | integer | Filing year. |
documentTitle | text | Document type (e.g. Form 40-F annual filing, Form 6-K filing, issuer report). |
documentPeriod | text | Filing date or reporting period. |
signalScore | integer | Sum of mentionsScore + governanceScore. Retained for backward compatibility with the original published dataset; new analyses should use the two axis columns below. |
aiTerms | text | Matched AI vocabulary with per-term counts (combined mentions + governance), e.g. artificial intelligence (2); ai governance (1). |
riskTerms | text | Adjacent-controls vocabulary (privacy, cybersecurity, vendor risk, etc.). Broader than AI governance; for context only. |
sourceType | text | Origin of the document (SEC 40-F, SEC 6-K/20-F, issuer PDF/HTML/text). |
url | url | Direct link to the source document for verification. |
mentionsScore | integer | Count of generic AI vocabulary matches (artificial intelligence, machine learning, generative AI, large language models, natural language processing, computer vision, deep learning, neural network, chatbot, recommendation system, plus product / vendor mentions). The AI Mentions axis. |
governanceScore | integer | Count of explicit AI governance vocabulary matches (AI governance, responsible AI, AI ethics, AI risk / system / impact assessment, AI lifecycle, AI risk management framework, ISO 42001, EU AI Act). Plus anchored "model risk management" and "human oversight" matches (require an AI-context token within ±400 / ±200 chars). The AI Governance axis. |
mentionsTerms | text | Distinct mention-axis labels that fired in this document, pipe-delimited with counts. |
governanceTerms | text | Distinct governance-axis labels that fired in this document, pipe-delimited with counts. |
lensSignals | text | Strict AI-linked disclosure-lens signals that fired in this document, including lens name, evidence tier, and top matched terms. Newly refreshed filing rows are scored from extracted filing text before compact public excerpts are emitted. A theme must appear near AI language. Lenses cover customer privacy, data sovereignty, compute sovereignty, open-weight/local models, third-party AI vendors, and AI cyber/abuse risk. |
broadLensSignals | text | Broader disclosure-lens signals using the same theme dictionaries without requiring AI proximity. Newly refreshed rows use extracted filing text where available, while legacy rows fall back to stored evidence snippets until refreshed. Useful for comparing broad topic disclosure against stricter AI-linked evidence. |
primaryPosture | text | Structured context for a lens match: internal deployment, customer product/service, investment exposure, risk disclosure, planned/exploratory use, denied/non-use, or unclear. |
publicationStatus | text | reviewed, screened, or candidate. Candidate evidence is kept out of the public research bundle. |
supportsUseClaim | boolean | True only when internal-deployment evidence has been manually reviewed. Automated screening never sets this flag. |
lensPostures | text | Compact per-lens publication status and posture in the public CSV. |
JSON structure (research-edition.json)
| Key | Description |
|---|---|
asOf | Release date of the dataset. |
canada.universe | Full TSX/TSXV listed-company counts (the broad-market denominator). |
canada.officialSourceCorpus | Parsed official-source documents and the issuers/companies they cover. |
canada.secFiledCorpus | SEC Form 40-F annual filings scanned. |
canada.eftsCorpus | SEC 6-K / 20-F filings located via EDGAR full-text search. |
canada.exchangeCoverage | Companies with AI evidence, by exchange. |
signalFilings[] | Per-document signal rows (the CSV is a flattened view of this), each with source URL and page/passage-level evidence. |
signalFilings[].lensSignals | Per-document AI-linked disclosure-lens matches with score, tier, confidence, review status, source context, terms, and short evidence excerpts. |
signalFilings[].broadLensSignals | Per-document broad disclosure-lens matches with score, tier, and terms. These are topic-disclosure signals, not AI-proximity evidence. |
topCompanies[] | Issuers ranked by total AI disclosure signal score, with compact per-company strict and broad lens summaries. |
lensSummary | Definitions, broad-disclosure counts, AI-linked evidence counts, document counts, and top companies for each disclosure lens. |
meta.publicPolicy | Defines the public research, professional pilot, private operating, and paid-customer licensing boundaries. |
meta.accuracy | Aggregate precision/recall evaluation state. Metrics are not publication-grade until every lens meets the minimum labeled sample size. |
yearSummary[] | Documents and signalled documents per year. |
listedUniverse[] | Every TSX/TSXV entity in scope. |
Scope & limitations
The listed directory covers all TSX/TSXV entities. Weekly metadata-delta checks identify new SEDAR+ continuous-disclosure filings through the current POC source, while SEC and issuer-hosted artifacts provide supplemental evidence. Retrieval failures are recorded separately from clean scans. Public classifications are reviewed or explicitly labeled automated screening and are separated by posture. Disclosure does not imply adoption, and only manually reviewed internal-deployment evidence may support a company-use claim.
Contact: info@airiskmanagement.ca