Jump to content

Research:Vital Knowledge Operationalization

From Meta, a Wikimedia project coordination wiki
This is an archived version of this page, as edited by Silva Selva (talk | contribs) at 17:45, 12 January 2026 (created research page with the second part of the VK research). It may differ significantly from the current version.
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)
This page documents a completed research project.


This project operationalized the Vital Knowledge research into concrete pipelines with four Wikimedia communities — Punjabi, Telugu, Uganda, and Singaporean Wikimedians. Through co-creation, we developed reusable methodologies for generating prioritized Vital Knowledge lists, validating the hypothesis that communities can operationalize VK research into actionable workflows.

Background

Previous research revealed that communities define "vital knowledge" differently based on local contexts, but all face challenges in systematically identifying and prioritizing content gaps. This project tested whether communities could co-create operational workflows using their preferred data sources (pageviews, education curricula, external encyclopedias) and validation processes (community consultations) to generate prioritized article lists.

Methods

We worked collaboratively with each community from October-December 2025 to:

  1. Co-design pipelines tailored to community definitions of "vital"
  2. Implement data gathering using community-preferred sources
  3. Generate draft prioritized lists through hybrid automated/manual approaches
  4. Validate through community endorsement processes

Each pipeline balanced three data points:

    • Pageviews and reader traffic patterns
    • Comparisons with other online encyclopedias
    • External reader interest signals (news trends, Google searches)

Results

Target Achievement

    • Baseline: 0 operational VK pipelines at project start
    • Target: 2+ communities with functional VK pipelines
    • Actual: 3 communities with a methodology mapping

Deliverables

  1. Data pipelines for Luganda, Telugu, and Punjabi Wikipedias to retrieve pageviews and prioritize articles
  2. Methodology framework for identifying gaps through comparison with existing online encyclopedias (Infopedia, Punjabipedia) - *documented as a replicable approach but not yet implemented*
  3. Reusable micro-task generator that can turn any list of Wikipedia articles into structured newcomer tasks
  4. Community draft lists:
    • Punjabi Wikimedians: 71 priority articles
    • Uganda: "Top 10 articles by category" list
    • Telugu Wikipedia: ~1,000 prioritized vital articles based on pageviews

Key Insights

  • Vital ≠ Popular: Telugu's case demonstrated zero overlap between their Vital Knowledge list and top-viewed Google pages, highlighting the tension between reader demand (contemporary pop culture) and encyclopedic importance (foundational, academic topics)
  • Hybrid approaches are necessary: No single data signal (pageviews, searches, curricula) proved sufficient alone
  • Language ≠ Culture: Single-language wikis often span multiple cultural contexts requiring regional adaptation
  • Consistent useful data points emerged: Despite different definitions of "vital", all communities found pageviews, external encyclopedia comparisons, and reader interest signals valuable

Operational Pipeline Model

Based on our co-creation work, we identified a reusable framework for VK list creation:

Phase 1: Foundation & Partnership

  • Community engagement and definition of "vital"
  • Identification of preferred data sources and validation processes

Phase 2: Data Gathering & Processing

  • Automated data collection (pageviews, cross-wiki comparisons)
  • Manual enhancement (community consultations, unwritten knowledge mapping)

Phase 3: List Creation & Prioritization

  • Gap analysis and triangulation across multiple data sources
  • Priority setting balancing community definitions of importance

Phase 4: Community Endorsement

  • Structured community feedback and consultation
  • Formal sign-off and ratification of final lists

Phase 5: Activation & Scaling

  • Conversion of lists into actionable tasks
  • Model documentation and adaptation for other communities

Challenges & Opportunities

Current Limitations

  • Tool usability: Pageview pipelines remain technical and not easily accessible for non-technical community members
  • Colonial language complexity: Defining "vital" in widely spoken colonial languages (e.g., Spanish, English) is highly region-dependent
  • Mapping unwritten knowledge: Still requires significant manual effort as automated signals (Google searches) surface popular rather than vital topics
  • Pipelines are not yet fully automated or scalable: Each requires community-dependent manual work

Identified Tool Development Pathways

1. Improved pageviews system - Communities consistently care about reader demand, however we should be mindful of the fact that pageviews is not the same as impact, since increasingly Wikipedia facts are being ingested via AI summaries 2. Structured search gap data partnerships - Better integration with platforms like Google 3. Adaptation of breaking news tools - The Enterprise team's open-access tool (https://phabricator.wikimedia.org/T412065) could be adapted for multicultural gap analysis 4. Modern cross-wiki linking tools - Needed to better identify missing interlanguage links

Conclusion

The hypothesis was supported: we successfully co-created Vital Knowledge pipelines with four communities, generating draft lists and reusable methods. While each community has a different definition of "vital," we identified consistent data points useful across all groups, confirming that a reusable, adaptable model for creating VK lists is possible.

The groundwork is now laid for activation. The next phase should focus on:

    • Activating communities to create content based on these validated lists
    • Experimenting with interventions to support quality contributions
    • Understanding whether prioritized vital knowledge lists can make content work less overwhelming and more engaging for small communities
    • Developing the identified tool improvements to increase pipeline scalability

See also

Diff Blog Posts