Jump to content

Next 25/Making Wikimedia's human collaboration visible

From Meta, a Wikimedia project coordination wiki
Making Wikimedia's human collaboration visible
A Next25 evidence and consent-building workspace
Status Exploratory workspace. No statement, feature, metric, or tool described here has been endorsed by the Wikimedia movement.
Core question How should Wikimedia help readers understand the people, communities, and collaborative processes behind its knowledge?
Purpose Gather evidence, make perspectives and objections visible, craft statements that may obtain consent, and identify responsible next steps.
Parent initiative Next25
Facilitator Schiste
Discussion Talk page
Evidence last reviewed 17 August 2026
First consent review To be scheduled

This page explores whether, why, and how Wikimedia should make the human and collaborative production of its knowledge more visible.

The page must remain neutral between possible outcomes, including the conclusion that no reader-facing change should be made.

Core question

[edit]

How should Wikimedia help readers understand the people, communities, and collaborative processes behind its knowledge?

Supporting questions include:

  • What do readers currently understand about how Wikimedia content is produced, revised, debated, and maintained?
  • Would greater visibility of collaboration improve understanding, trust calibration, recognition, or participation?
  • Which forms of contribution can be represented fairly, and which would be erased by common metrics?
  • What privacy, safety, equity, accessibility, performance, and gaming risks would new interfaces create?
  • What should be shown collectively, what may be shown about individuals, and what should not be shown?
  • Which decisions belong to local communities, tool maintainers, product teams, the Wikimedia Foundation, or other bodies?
  • What small and reversible experiments could generate better evidence?

First discussion round

[edit]

The first round is about positioning two most fundamental statements:

P1: The human and collaborative production of Wikimedia knowledge is part of Wikimedia's public value.

P2: Wikimedia should help readers understand that its knowledge is created, revised, debated, and maintained collaboratively by humans.

Subsequent rounds will aim to drill down on implementation rules, constraints, etc...

Scope

[edit]

In scope

[edit]
  • Reader understanding of Wikimedia's collaborative production.
  • Article history, revision, and provenance interfaces.
  • The relationship between Wikimedia content and external reuse, including reuse by search, social, and AI systems.

How to participate

[edit]

Participants can contribute in five ways:

  1. Record a position. Use the consent system format in the talk page.
  2. Raise a material objection. Describe the plausible harm or constraint and propose an amendment, safeguard, alternative, or stop condition where possible.
  3. Add evidence and examples Include the method, population, finding, limitations, and relevance, not only a link.
  4. Improve a statement. Propose exact replacement wording and explain which concern it addresses.
  5. Propose or assess an experiment. State the hypothesis, affected people, decision owner, success measures, risks, and reversibility.

Use this page for the current synthesis and the talk page for debate. Important arguments should be summarised here so that the outcome does not depend on reading every discussion thread.

Please sign talk-page contributions with ~~~~.

Research, interesting reading and past experiments

[edit]

This is a starting map, not a complete literature review.

Source Question and method Findings relevant here Limitations and cautions
Elmimouni, Forte and Morgan, Why People Trust Wikipedia Articles (2022)[1] Survey of 1,716 English Wikipedia readers in 105 countries, followed by 17 interviews, examining credibility assessment. Readers used direct experience, visible credibility indicators, and mental models of Wikipedia's editorial process. English Wikipedia recruitment and a small interview sample limit generalisation. The study does not directly test contributor displays.
Flöck and Acosta, WikiWho (2014)[2] Development and evaluation of token-level attribution for revisioned wiki text. Demonstrates that surviving text can be computationally traced to revisions with high reported precision on the study's gold standard. Token provenance is not a normative definition of authorship, effort, quality, responsibility, or value. Results depend on revision data and algorithmic choices.
Lucassen and Schraagen, evaluation of WikiTrust (2011)[3] Eye-tracking experiment evaluating colour overlays intended to communicate text trust. Heavy visual signalling made reading more difficult, reduced trust for heavily coloured articles, and received low usefulness ratings. A historical prototype with a specific visual treatment. It does not show that all provenance or trust interfaces are harmful.
Recognition on the wikis (Wikimedia Foundation, completed 2026) Qualitative study with 23 participants across English, Spanish, and French Wikipedias on recognition experiences and concepts. Simple appreciation was valued; personalised and group or project-based recognition received stronger enthusiasm. Optionality was central. Light-touch recognition alone did not strongly motivate most participants. Small qualitative sample in three language communities. Recognition by peers is not equivalent to reader-facing attribution.
Attribution Framework: Ecosystem Growth research support (Wikimedia Foundation, completed 2026) Survey of 1,022 respondents in the United States and nine moderated sessions testing off-platform attribution and calls to action. Messages emphasising Wikipedia's value, collaborative nature, or free-knowledge identity generally outperformed neutral language. Perceived relevance and alignment with Wikipedia identity predicted engagement intent. United States, English-language, off-platform context. This was not a direct test of article-page contributor names or rankings.
Restivo and van de Rijt, informal rewards experiment (2012)[4] Randomised experiment awarding barnstars to a subset of highly productive, previously unrecognised Wikipedia contributors. The treated group increased productivity and was more likely to receive later peer recognition. The intervention involved peer recognition of a selected group of productive editors. It does not establish that public reader-facing rankings motivate contributors generally.
Grounding Data Provenance in Wikipedia Editors' Expectations of Attribution (research begun 2026) Ongoing research into editors' expectations when Wikimedia content is reused by large language models and into possible provenance mechanisms. No final findings should be inferred while the study is in progress. Ongoing work. Results are not yet available for statement validation.
Why readers trust Wikipedia: bibliography Curated bibliography covering reader trust, credibility assessment, and related research. Provides a broader evidence base for evaluating trust claims. Individual sources vary in method, date, project, and relevance. Each must be assessed independently.
Roles of contributors Research and synthesis on the varied roles contributors perform. Supports examining contribution beyond the production of surviving article prose. Role taxonomies do not automatically provide a fair or computable recognition metric.

Previous Wikimedia work and discussions

[edit]
Work or discussion What it does Relevance Current caution
Phabricator T2639: annotate or blame-like authorship features Long-running proposals to expose who introduced current text and improve attribution in article history. Documents technical ideas, community interest, and unresolved design questions. A task history is not current community consent, and “blame” or “authorship” terminology can mislead.
Who Wrote That? Browser extension that lets readers inspect who introduced text and the revision in which it appeared. Demonstrates an interactive, on-demand approach rather than a permanent ranking. Depends on attribution data and exposes only what its method can trace.
XTools Authorship Estimates current-text authorship by character count, using WikiWho data. Existing tool for comparing current surviving text associated with editors. Character count excluding spaces is a metric choice, not a complete account of authorship or contribution.
WikiWho and WhoColor Research and tools for tracing text tokens to revisions and visualising authorship, age, and conflict. Technical foundation for some current-text provenance experiments. Text survival can overvalue bulk prose and undervalue research, correction, maintenance, discussion, and non-text work.
RevisionSlider Adds a visual control for navigating between revisions in the diff interface. Example of making revision history easier to explore without assigning article ownership. Primarily serves history exploration rather than contributor recognition.
WikiTrust Historical research and prototypes estimating text reputation and presenting visual trust signals. Provides evidence about both the appeal and risks of visible article-level signals. Trust scores and colour overlays can burden readers and create false precision.
Wikimedia Attribution Framework Recommends clearer and more consistent recognition of Wikimedia as the source of externally reused content, including signals of human collaboration and pathways back to Wikimedia. Connects source attribution, Wikimedia identity, human contribution, and participation. Primarily concerns external reuse. It does not determine how article pages should rank or name contributors.
Attribution API Provides structured data to help reusers generate fair and consistent attribution. Possible infrastructure for external provenance and attribution experiences. API fields should not be treated as a definitive measure of contributor value.
Recognition on the wikis Examines how contributors experience and understand recognition. Provides recent qualitative evidence for optional, contextual, and community-based approaches. Recognition preferences vary. A single global interface is unlikely to fit every context.

Please add links to relevant Village Pump discussions, Requests for Comment, Community Wishlist proposals, research pages, product studies, Phabricator tasks, gadgets, tools, and local community experiments.

Affected perspectives

[edit]

A credible consent process must seek affected perspectives, not only general opinions.

Perspective Possible value Questions and risks to surface
Readers unfamiliar with editing Better understanding of collaboration, revision, and participation. Will the interface clarify the process or create clutter, false authority, or confusion?
Experienced editors Recognition of sustained work and clearer links to contribution history. Will the metric erase maintenance, sourcing, discussion, or collaborative refinement?
New and occasional editors A visible signal that contribution is possible and noticed. Will rankings make contribution feel closed, competitive, or dominated by established editors?
Unregistered and temporary-account contributors Recognition of anonymous participation where appropriate. How will masking, identifier changes, privacy, and unwanted exposure be handled?
Contributors to sensitive topics Potential acknowledgement of difficult and valuable work. Could visibility increase harassment, legal exposure, surveillance, or pressure?
Small-language and under-resourced projects Greater visibility for communities whose work is often overlooked. Will data coverage, interface quality, and research validation be weaker than for large Wikipedias?
Patrollers, administrators, mediators, and anti-abuse contributors Recognition of work that protects content and communities. Most of this work is not represented by surviving article text.
Translators, template and module editors, bot operators, Commons contributors, and Wikidata contributors Recognition of dependencies and non-prose work behind an article. Can a reader-facing interface represent these contributions without becoming incomprehensible?
WikiProjects, campaigns, affiliates, and organisers Visibility for collective production, mentorship, and coordinated work. How should collective recognition relate to individual credit and local community norms?
Product, design, research, engineering, and data teams Testable hypotheses and reusable provenance infrastructure. Data correctness, performance, maintainability, accessibility, and opportunity cost.
Privacy, Trust and Safety, security, and legal specialists Better-defined safeguards and responsible reuse. Suppression, masking, harassment, data protection, legal attribution, and abuse cases.
External reusers Clearer ways to acknowledge Wikimedia and direct people to sources or participation. Will attribution remain accurate, comprehensible, and implementable across contexts?
People using assistive technologies or low-bandwidth devices Better access to provenance where designed accessibly. Extra labels, colour, interactions, scripts, and data requests may create barriers.

Hypotheses to consider

[edit]

These are hypotheses, not established facts.

Hypothesis Evidence that would support it Evidence that would challenge it
A clearer collaboration signal improves readers' understanding of how Wikimedia content is produced.
Provenance information improves trust calibration.
Contextual recognition improves contributor experience.
Visible collaboration creates a useful participation pathway.
Collective attribution communicates Wikimedia's human distinctiveness in external reuse.
Shared provenance infrastructure enables better community-built tools.

Benefits to investigate

[edit]

Potential benefits include:

  • helping readers understand that Wikimedia content is built, revised, and maintained by people;
  • making collaboration and revision easier to inspect;
  • connecting readers to histories, communities, and participation pathways;
  • recognising individual and collective effort where contributors consider that recognition meaningful;
  • improving external attribution and source visibility;
  • supporting research and community-built provenance tools;
  • distinguishing Wikimedia's open and participatory process from knowledge products that obscure their sources.

None of these benefits should be assumed. Each needs evidence, affected-perspective review, and a testable hypothesis.

Open objections and risks

[edit]

The register below begins with foreseeable objections. It should record changes over time rather than deleting objections once wording is amended.

Material objection Response required before proceeding Mitigation implemented
Visible-text metrics can reward bulk prose while erasing research, sourcing, correction, translation, media work, structured data, templates, discussion, mentoring, patrolling, and organising. State what is measured, avoid totalising language, test non-ranking and collective alternatives, and document omitted work.
Greater visibility may expose contributors working on sensitive topics to harassment, surveillance, legal pressure, or unwanted attention. Privacy and safety review, data minimisation, treatment of masking and suppression, opt-out or non-display paths where appropriate, and explicit stop conditions.
Surviving text is not equivalent to quality, correctness, expertise, care, responsibility, or endorsement. Use precise language, test reader interpretation, and never label a metric as a quality or authority score.
Bots, templates, transclusions, imports, translations, page moves, merges, deletions, and cross-project dependencies make attribution incomplete or misleading. Publish handling rules, known error cases, project coverage, and confidence or uncertainty.
Named or ranked contributors may be interpreted as article owners, official representatives, experts, or people responsible for the current version. Test language and comprehension, avoid ownership framing, and provide direct explanation of collaborative revision.
Additional messages, colours, names, and controls may create clutter, cognitive burden, accessibility barriers, or slower pages. Compare on-demand and low-salience designs, conduct accessibility review, measure performance, and include a no-change control.
Tools and studies may work best on large Wikipedias, increasing inequity and making smaller projects appear less active or less legitimate. Include multilingual pilots, document data gaps, avoid cross-project comparisons, and give local communities authority over adoption.
Public metrics and rankings can be gamed and may redirect effort from useful work towards visible score optimisation. Threat modelling, gaming tests, monitoring, reversible deployment, and avoidance of competitive leaderboards.
Historical identifiers may reveal names or links that contributors no longer expect to be foregrounded, even when the underlying history remains public. Distinguish data availability from responsible prominence, review edge cases, and minimise unnecessary amplification.
A pilot can create real harm even when described as experimental, especially if deployed at scale before safeguards are validated. Begin with research prototypes or opt-in tests, limit exposure, predefine rollback, and monitor harm from the first release.

Experiment register

[edit]

Wikipeople - highlighting the human processus and contributors

[edit]
Tool fr:Utilisateur:Schiste/wikipeople
Maintainer Schiste
Status Exploratory prototype
Relation to this discussion One implementation of experiment to showcase human editors and improve history page to convert curious readers
What it demonstrates A possible way to surface contributors associated with an article through a defined metric.
What it does not establish That its ranking is a complete or neutral account of authorship, effort, care, value, expertise, responsibility, or endorsement.

Questions for the Wikipeople prototype

[edit]
  • Does the prototype help readers understand collaboration, or mainly create a list of prominent names?
  • Do highlighted contributors recognise the result as fair and meaningful?
  • Who is systematically omitted?
  • How often does the ranking change after small edits?
  • Can the ranking be manipulated cheaply?
  • Does it create incentives to preserve one's text, make large edits, or avoid collaborative rewriting?
  • Does it foreground people who do not want additional prominence?
  • Can it explain errors and limitations without overwhelming readers?
  • Does it work comparably across projects and languages?
  • Is a named ranking more effective than a collective count, an unranked list, a role summary, or an improved history link?
  • What evidence would justify continuing, changing, or stopping the prototype?

References

[edit]
  1. Houda Elmimouni, Andrea Forte and Jonathan T. Morgan, “Why People Trust Wikipedia Articles: Credibility Assessment Strategies Used by Readers”, OpenSym 2022, DOI: 10.1145/3555051.3555052.
  2. Fabian Flöck and Maribel Acosta, “WikiWho: Precise and Efficient Attribution of Authorship of Revisioned Content”, WWW 2014, DOI: 10.1145/2566486.2568026.
  3. Teun Lucassen and Jan Maarten Schraagen, “Evaluating WikiTrust: A trust support tool for Wikipedia”, First Monday, volume 16, number 5, 2011, DOI: 10.5210/fm.v16i5.3070.
  4. Michael Restivo and Arnout van de Rijt, “Experimental Study of Informal Rewards in Peer Production”, PLOS ONE 7(3): e34358, 2012, DOI: 10.1371/journal.pone.0034358.

Further reading

[edit]

See also

[edit]