huby logo
huby logo
Version 1.1 — August 2026

huby AI Product
Evaluation Methodology

Trusted AI, independently evaluated

1

Purpose and Scope

This document describes the methodology huby uses to evaluate AI products and produce its Product Transparency & Evaluation Reports. It is published in full to allow readers, product owners, and enterprise buyers to independently assess the validity and limitations of our findings.

This methodology applies to all AI products evaluated on the huby platform. Product-type-specific adaptations are documented in companion framework specifications where applicable.

This document is the authoritative reference for how products are scored. For a step-by-step view of how an evaluation moves from submission through to publication, see the evaluation process. Where the two differ, this document takes precedence.

2

Product Classification

Every evaluation begins with product classification. We first assign the product to a product type (e.g. software engineering, creative toolkit, AI Agent) and then identify a product subtype (e.g. coding assistant, multi-modal, knowledge management) by mapping it against the competitive landscape. This classification determines which subcategories and atomic factors are activated within the standard 6 categories of huby's evaluation framework.

3

Evaluation Framework

3.1

Structure

The evaluation framework is organized in three tiers:

TierDescriptionExample
CategorySix broad dimensions common across all product typesSecurity
SubcategoryMeasurable dimensions underlying a particular category Access Control Mechanisms
Atomic factorThe specific, verifiable data point or capability that underlies a specific subcategoryAvailability of native multi-factor authentication

Structural variability by design. The number of subcategories within each category, and the number of atomic factors within each subcategory, are not uniform. They are determined during framework establishment based on the complexity and relevance of each dimension for the specific product subtype being evaluated. A category with broader regulatory surface area (such as Privacy) may carry more subcategories than a category where the product subtype has a narrower scope. Similarly, a subcategory covering a complex domain (such as infrastructure security) will carry more atomic factors than one covering a more bounded capability. This variability is intentional — it ensures the framework reflects the actual shape of the evaluation problem rather than imposing artificial symmetry. The complete subcategory and factor inventory for each product subtype is documented in the corresponding framework specification.

3.2

The Six Evaluation Categories

Quality

The degree to which the product delivers accurate, consistent, and reliable outputs relative to its stated purpose. Indicators include output accuracy and consistency, user experience design, product documentation and developer support, independent benchmarks, and reported defect rates and resolution times.

Use Cases & Pricing

The breadth and depth of supported use cases, the presence of novel or differentiated capabilities, and integration with relevant ecosystems. Pricing assessment covers tier structure, free access provisions, licensing model transparency, and value for money relative to competitive alternatives. Adoption signals such as user growth and retention are considered where data is available.

Security

The controls in place to protect the product, its infrastructure, and user data from unauthorized access and exploitation. Evidence considered includes third-party security audit reports, published CVEs and remediation records, penetration testing outcomes, static and dynamic code analysis results, API security posture, data encryption in transit and at rest, and access control mechanisms including multi-factor authentication and role-based access where applicable.

Privacy

The transparency, consistency, and user-centricity of the product's data handling practices. Assessment covers privacy policy clarity and enforceability, relevant compliance certifications (SOC 2, GDPR, CCPA/CPRA, HIPAA where applicable), data sharing practices with third parties, data retention and deletion controls, and user rights regarding access, portability, and consent.

Sustainability & Reliability

Evaluated across two dimensions. Company sustainability assesses organizational longevity indicators including funding status and runway, revenue trajectory, team depth, and market position. Service reliability assesses uptime commitments, published SLAs, historical incident records, and performance benchmarks. Environmental and social sustainability practices are noted where publicly documented.

Impact, Ethics & Safety

The degree to which the product and its developer act responsibly toward users and society. Assessment covers the existence and substance of a published ethics policy, demonstrated practices around bias mitigation and model transparency, controls to prevent misuse and harm to users, societal applications and misuse risk, and the company's track record of responding to ethical failures or safety incidents.

4

Scoring

4.1

Scale

All scores are reported on a 1.0–5.0 scale at atomic factor, subcategory, and category levels.
4.2

Score Definitions

5.0

Exemplary

Strong, independently verified evidence of best-in-class practice. Formal audits, third-party certifications, or peer-reviewed assessments confirm performance.

4.0–4.9

Strong

Multiple credible, corroborated sources substantiate the claim. Minor gaps exist but do not materially undermine the finding.

3.0–3.9

Adequate

Evidence is present but limited, mixed, or partially corroborated. The product meets a basic standard but with notable gaps or inconsistencies.

2.0–2.9

Weak

Evidence suggests meaningful shortfalls. Claims made by the product owner are not substantiated by independent sources, or documented failures are present.

1.0–1.9

Deficient

Material failures or absence of the capability/control. Significant risk to users or enterprise buyers exists.

Scores between integers reflect graduated evidence quality within a band (e.g., 3.7 indicates the upper range of "Adequate" but below "Strong" threshold). We do not assign scores that are not corroborated by the evidence record for that factor.

4.3

Score Aggregation

Scores are aggregated upward through the three-tier evaluation hierarchy using a weighted model established during framework design, not a simple average. Weights are assigned at each level — across atomic factors within a subcategory, and across subcategories within a category — before any data collection begins. This ensures weights reflect considered analytical judgment about relative importance rather than being adjusted post-analysis to fit observed results.

Atomic factor to subcategory. Each atomic factor within a subcategory carries a defined weight that reflects its relative criticality for the product subtype. For example, within an Access Control subcategory, the presence of multi-factor authentication may be weighted more heavily than audit log granularity, because its absence represents a more material risk to end users. Factor weights within each subcategory sum to 100%.

Subcategory to category. Each subcategory within a category carries a defined weight reflecting its centrality to the category's overall assessment. Subcategory weights within each category sum to 100%.

Weights are defined in the framework specification for each product subcategory. They are not adjusted after data collection begins. Any post-analysis weight change would require a full re-evaluation and a new report version for the entire cohort of AI products.

N/A factors (those excluded due to insufficient evidence) are removed from the weight distribution, and the remaining factor weights are normalized to sum to 100% before aggregation. This prevents evidence gaps from artificially deflating scores.

Category scores are further aggregated into a single composite score based on individual category scores and their respective weightage. Users can easily see assigned scores at different levels.

4.4

Sample Score Breakdown — AI Chatbot

The following illustrates how pillar and subcategory scores appear in a huby evaluation report for an AI chatbot product. Each category shows its overall pillar score (orange) alongside the individual subcategory scores (tan) that roll up into it.

Pillar score
Subcategory score
1.0Quality
4.3
Multimodal Reasoning Reliability
4.1
Creative Generation Quality
4.5
Agentic Workflow Reliability
4.3
User Experience & Accessibility
4.3
Developer Ecosystem Maturity
4.3
2.0Privacy
4.4
Inference Data Governance
4.3
Enterprise Privacy Controls
4.5
3.0Security
4.2
Model & Platform Security
4.2
Enterprise Governance Security
4.3
4.0Sustainability & Reliability
4.3
Operational Reliability
4.2
Business Sustainability
4.7
5.0Use Cases & Pricing
4.5
Enterprise Workflow Coverage
4.5
Pricing & Economic Efficiency
4.4
6.0Impact, Ethics & Safety
4.0
AI Safety Governance
4.2
Copyright & Attribution Governance
3.9
Psychological & Societal Impact
3.5
5

Data Collection

5.1

Source Hierarchy

huby uses a tiered source framework. Higher-tier sources take precedence when sources conflict. Every claim that affects a score must be traceable. Secondary sources (Tiers 2–5) are cited with a publicly accessible URL. Tier 1 findings are our own measurements and have no external URL, so they are instead documented with the test date, the product version and access tier tested, and the sample size.

TierSource TypeExamples
Tier 1Primary research and testing by hubyIdentical tests run across every product in a cohort, against the framework for that category; test date, product version, access tier and sample size recorded
Tier 2Independent third-party audits and certificationsSOC 2 reports, CVE databases, named security firm assessments, peer-reviewed publicly available research
Tier 3Product owner documentationOfficial privacy policies, terms of service, technical documentation, changelogs
Tier 4Structured user and practitioner evidenceEnterprise case studies from named customers, verified professional community posts
Tier 5Anecdotal community evidenceReddit threads, social media posts, informal reviews
5.2

huby's Research and Testing Data

When we decide to research a particular category of AI products, we identify a set of cohorts in that category that broadly speaking fall in that category. Then for each of our standard six evaluation categories we establish a common set of evaluation subcategories and factors. We also establish a common test criteria for them.

We then run those tests ourselves. We test whatever access is available to us: the publicly available free or trial tier, or a test account or API key provided by the product owner. Where a product owner provides access, the access carries no conditions — no review of findings before publication, no embargo, and no influence over scoring. This research and data then becomes part of the larger data set we use for evaluation and scoring.

Every test is run identically across all products in the cohort, on the same inputs, within as narrow a time window as practical. Each report records the test date, the product version tested, the access tier used, and the number of test items, so that a later re-evaluation can distinguish a genuine change in the product from a change in how we measured it.

5.3

Product Owner Data

Since many AI products don't have a long history, we require less known product owners to submit their product documentation, technical specifications, and supporting materials directly to huby for consideration. Product owner submissions do not receive preferential weighting.

huby is independent and does not charge product owners for evaluation, for listing, or for either report. There is no paid tier, no expedited review, and no way for a product owner to pay for placement, for a score, or for a score to be reconsidered. Evaluations are not contingent on a product owner engaging with us at all — many are initiated by huby without the owner's involvement. Should any commercial relationship with a product owner exist in future, it will be disclosed in the relevant report and this section will be updated first.

5.4

Competitive Context

Where sufficient public data exists, scores are contextualized against the assessed product's competitive set. This does not change the score itself but provides readers with relative positioning. Where competitive benchmarking is not possible, this limitation is disclosed.

5.5

Data Normalization

Raw evidence is normalized across four dimensions before scoring: recency (evidence older than 24 months is discounted unless no newer data exists), source independence (per the hierarchy above), and specificity (general claims are weighted below specific, verifiable assertions).

6

Analytical Process

6.1

Independence and Conflict Management

Our analysis is not influenced by a product owner. We fully disclose any work that we do with product owners. Score assignments at subcategory and above go through multiple reviews before they are published.

Who controls a product's visibility on huby, and what that does not control. Where a product owner submits their own product, they set two flags: one confirming the product is ready for evaluation, and one enabling the discovery of the product to go live on huby platform. They do not govern the evaluation, the scores, or the written assessment, none of which are shared with the owner for approval before publication. Where huby initiates an evaluation — which is the case for most established products, and for every product in a category-wide comparative study — huby sets both flags and publishes regardless of the outcome.

6.4

Report Versioning and Updates

Each report carries a production date and is reviewed for material updates on a regular basis. A revision may be brought forward when we are notified of a significant product change, a security incident, or a regulatory development that warrants earlier review.

7

Report Structure

huby produces two report types for each evaluated product:

Public Transparency Report

Covers all six evaluation categories with subcategory-level scores and written assessments. Evidence is cited with verifiable URLs. This report is available to all users of the huby platform at no cost.

Detailed Owner Report

Produced for the product owner. Includes atomic factor-level scores and analysis, specific improvement recommendations mapped to each factor, and competitive gap analysis. This report is confidential to the product owner offering them an opportunity to improvise their product.

8

Limitations

Users of huby reports should be aware of the following inherent limitations:

  • Reports are based on our testing, publicly available evidence and product owner submissions at the time of evaluation. They do not constitute a formal security audit, legal compliance certification, or financial advisory opinion.
  • huby's own tests are conducted on a defined sample of test items, on a specific product version and access tier, at a point in time. They are not exhaustive. Results drawn from small samples carry uncertainty and indicate direction rather than a precise ranking.
  • Scores reflect the state of the product at the report production date and may not reflect subsequent changes.
  • The absence of evidence for a capability or control does not confirm its absence — it reflects the limits of publicly available information at evaluation time.
  • At huby we are not lawyers, certified security auditors, or financial advisors. Where legal, security, or financial interpretation is required, readers should seek qualified professional advice.
9

Terminology

TermDefinition
Atomic factorThe smallest unit of evaluation; a specific, verifiable capability or control
CorroborationConfirmation of a claim by an independent source separate from the originating source
Coverage failureA case where a product cannot process a test input — through a length limit, quota, or error. Scored as a result, not recorded as an evidence gap
Evidence gapA topic within the framework where insufficient public evidence exists to assign a score
Primary testingTier 1 evidence; a measurement huby produced by running a defined test against the product itself, rather than a claim reported by another party
N/ANot Assessable; a factor excluded from score calculations due to an evidence gap
Product subtypeA refined classification of a product within a product type, used to determine applicable subcategories
RollupThe aggregation of atomic factor scores into subcategory scores, and subcategory scores into category scores

© 2026 huby. All rights reserved. This methodology document is published under huby's commitment to evaluation transparency. Reproduction for non-commercial reference is permitted with attribution.