Jump to content

Machine Learning/de: Difference between revisions

From mediawiki.org
Content deleted Content added
Created page with "Unser Team kümmert sich um die Entwicklung und Verwaltung von Machine-Learning-Modellen für Endnutzer sowie um die Infrastruktur, die für die Entwicklung, das Training und den Einsatz dieser Modelle erforderlich ist."
FuzzyBot (talk | contribs)
Updating to match new version of source page
 
(30 intermediate revisions by 2 users not shown)
Line 6: Line 6:
| group = Technology & Product
| group = Technology & Product
| lead = [[User:CAlbon (WMF)|Chris Albon]]
| lead = [[User:CAlbon (WMF)|Chris Albon]]
| team = [[User:AChou-WMF|Aiko Chou]], [[User:KBazira (WMF)|Kevin Bazira]], [[User:TKlausmann (WMF)|Tobias Klausmann]], [[User:LToscano (WMF)|Luca Toscano]], [[User:ISarantopoulos-WMF|Ilias Sarantopoulos]], [[Machine_Learning/almuni|(list of former staff and volunteers)]]
| team = [[User:AChou-WMF|Aiko Chou]], [[User:KBazira (WMF)|Kevin Bazira]], [[User:GKyziridis-WMF|Georgios Kyziridis]], [[User:TKlausmann (WMF)|Tobias Klausmann]], [[User:OKarakaya-WMF|Ozge Karakaya]], [[User:BWojtowicz-WMF|Bartosz Wójtowicz]], [[User:ISarantopoulos-WMF|Ilias Sarantopoulos]], [[User:DPogorzelski-WMF|Dawid Pogorzelski]], [[Machine_Learning/almuni|(list of former staff and volunteers)]]
| Phabricator = machine-learning-team
| Phabricator = machine-learning-team
| start = July 2017
}}
}}


Line 14: Line 15:
Unser Team kümmert sich um die Entwicklung und Verwaltung von Machine-Learning-Modellen für Endnutzer sowie um die Infrastruktur, die für die Entwicklung, das Training und den Einsatz dieser Modelle erforderlich ist.
Unser Team kümmert sich um die Entwicklung und Verwaltung von Machine-Learning-Modellen für Endnutzer sowie um die Infrastruktur, die für die Entwicklung, das Training und den Einsatz dieser Modelle erforderlich ist.


<span id="Current_projects"></span>
<div lang="en" dir="ltr" class="mw-content-ltr">
== Current projects ==
== Aktuelle Projekte ==
* [[m:Machine learning models|Modellkarten für maschinelles Lernen]]
</div>
* [[metawiki:Machine_learning_models|<span lang="en" dir="ltr" class="mw-content-ltr">Machine Learning Model Cards</span>]]
* <span lang="en" dir="ltr" class="mw-content-ltr">The [[Machine Learning/Modernization|Machine Learning Modernization]] Project</span>
* [[wikitech:Machine Learning/LiftWing|Lift Wing]] - Eine skalierbare Infrastruktur für maschinelles Lernen auf Kubernetes mit KServe.
* The [[Machine Learning/Modernization|<span lang="en" dir="ltr" class="mw-content-ltr">Machine Learning Modernization</span>]] <span lang="en" dir="ltr" class="mw-content-ltr">Project</span>
* [[wikitech:Machine_Learning/LiftWing|Lift Wing]] - <span lang="en" dir="ltr" class="mw-content-ltr">A scalable machine learning model serving infrastructure on Kubernetes using KServe.</span>


For archived projects, [[Machine_Learning/archived_projects|see this list]].
<span lang="en" dir="ltr" class="mw-content-ltr">For archived projects, [[Special:MyLanguage/Machine Learning/archived projects|see this list]].</span>


<span id="Contact_us"></span>
<div lang="en" dir="ltr" class="mw-content-ltr">
== Contact us ==
== Kontaktiere uns ==
Du hast eine Frage? Möchtest du mit dem Team oder unserer Gemeinschaft von Freiwilligen über maschinelles Lernen sprechen?
</div>
Hier sind die besten Möglichkeiten, mit uns in Kontakt zu treten.
<span lang="en" dir="ltr" class="mw-content-ltr">Have a question? Want to talk to the team or our community of volunteers about machine learning?</span>
<span lang="en" dir="ltr" class="mw-content-ltr">Here are the best ways to connect with us.</span>


{{ContentGrid
{{ContentGrid
Line 33: Line 32:


{{Colored box
{{Colored box
|title = <span lang="en" dir="ltr" class="mw-content-ltr">Team Chat</span>
|title = Team Chat
|content = <span lang="en" dir="ltr" class="mw-content-ltr">Discuss machine learning and watch the team work joining our public IRC chatroom {{irc|wikimedia-ml}} on <code>[https://libera.chat/guides/connect irc.libera.chat]</code>.</span>
|content = Diskutiere über maschinelles Lernen und beobachte das Team bei der Arbeit in unserem öffentlichen IRC-Chatroom {{irc|wikimedia-ml}} auf <code>[https://libera.chat/guides/connect irc.libera.chat]</code>.
}}
}}


{{Colored box
{{Colored box
|title = Aktive Arbeitspinnwand
|title = <span lang="en" dir="ltr" class="mw-content-ltr">Active Work Board</span>
|content = <span lang="en" dir="ltr" class="mw-content-ltr">Have a particular task you want to discuss or work on, join our public {{ll|Phabricator}} board.</span>
|content = Wenn du ein bestimmtes Anliegen hast, das du diskutieren oder bearbeiten möchtest, tritt unserem öffentlichen Board {{ll|Phabricator}} bei.
[[phab:project/view/1901/|<span lang="en" dir="ltr" class="mw-content-ltr">Visit our work board</span>]]
[[phab:project/view/1901/|Besuche unser Board]]
}}
}}


}}
}}


<span id="What&#039;s_new?"></span>
<div lang="en" dir="ltr" class="mw-content-ltr">
== What's new? ==
== Was ist neu? ==
* {{ymd|2024|9|20}}
</div>
** <span lang="en" dir="ltr" class="mw-content-ltr">We are continuing to integrate the article-country model into Liftwing.</span> <span lang="en" dir="ltr" class="mw-content-ltr">The article country model predicts which countries any particular model will be applicable for, and it's an extension of the article-topic model, which we have used for years.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We're trying different approaches to build vllm (a high-throughput and memory-efficient system designed for serving large language models) and ROCm (the code that allows the CPU to talk to AMD GPUs) with Ubuntu.</span> <span lang="en" dir="ltr" class="mw-content-ltr">This is part of the work of making production LLMs on Liftwing possible.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We're currently working on configuring the ML Lab servers. These are for model training.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Updated the rec-api image deployment model.</span> <span lang="en" dir="ltr" class="mw-content-ltr">Deployed the reference need model to production.</span>
* {{ymd|2024|8|23}}
** <span lang="en" dir="ltr" class="mw-content-ltr">Following up on recurring issue reported by the Structured Content team: The MediaDetection API can access the logo-detection endpoint via mwdebug1001.eqiad.wmnet and mwdebug2001.codfw.wmnet, but can't access it on k8s-mwdebug</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Adding logo-detection documentation to the API portal docs.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Investigating occasional slow queries on LiftWing when using some RevScoring models</span>
**<span lang="en" dir="ltr" class="mw-content-ltr">Continuing the remaining work on the pre-save revertrisk model.</span> <span lang="en" dir="ltr" class="mw-content-ltr">This model is designed to provide a vandalism prediction before an edit is saved to Wikipedia (and thus doesn't have a revision ID)</span>
**<span lang="en" dir="ltr" class="mw-content-ltr">Work continues on upgrading kserve to 0.13</span>
**<span lang="en" dir="ltr" class="mw-content-ltr">Initializing install config for the GPU hosts in eqiad</span>
* {{ymd|2024|6|14}}
** <span lang="en" dir="ltr" class="mw-content-ltr">Apologies on my end for the delay in updates, I got covid.</span>
* {{ymd|2024|5|17}}
** <span lang="en" dir="ltr" class="mw-content-ltr">Working continues on the Logo Detection model. We made an example logo-detection model-server that processes base64 image objects instead of image URLs and sent it to the Structured Content team for their thoughts.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Work continues on the HuggingFace model server.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">General bug fixes and improvements.</span>
* {{ymd|2024|5|10}}
* {{ymd|2024|5|10}}
** Die Arbeit am Logo-Erkennungsmodell geht weiter. Diese Woche haben wir im Team darüber diskutiert, ob das kodierte Bild direkt an LiftWing gesendet werden soll oder nicht. Alternativ würden wir eine URL mit dem Speicherort des Bildes erhalten, auf die LiftWing dann zugreifen und es herunterladen kann. Das ist wichtig, weil es sich auf die Größe der REST-Nutzdaten auswirkt, insbesondere bei gebündelten Anfragen.
** <span lang="en" dir="ltr" class="mw-content-ltr">Work continues on the Logo Detection model. The issues we discussed this week as a team whether or not the encoded image would be sent directly to LiftWing. Alternatively, we'd receive a URL of the image's location, which we would then allow LiftWing to access/download from. This matters because it affects the size of the REST payload, particularly with batched requests.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We are still working on the HuggingFace model server GPU issue (i.e. it won't recognize our AMD GPU). There are a number of possibilities as to why, but we want this resolved before we finalize our order for this fiscal year.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We are still working on the HuggingFace model server GPU issue (i.e. it won't recognize our AMD GPU). There are a number of possibilities as to why, but we want this resolved before we finalize our order for this fiscal year.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">A number of misc bug fixes and improvements.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">A number of misc bug fixes and improvements.</span>
* {{ymd|2024|5|03}}
* {{ymd|2024|5|03}}
** <span lang="en" dir="ltr" class="mw-content-ltr">Our big Istio refactoring is underway ([https://docs.google.com/presentation/d/1-C7_L51xJCStRWpX2mXOTluIC2CUO_6Umpna3lPL5Es/edit#slide=id.p slides])! This refactoring will allow us to remove a lot of networking logic out of individual model containers. For example, currently if there was some changing to the `discovery.wmnet` endpoint (WMF's internal endpoint for APIs), we'd have to update hundreds of individual model containers and redeploy them. This refactoring removes this need entirely.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Our big Istio refactoring is underway ([https://docs.google.com/presentation/d/1-C7_L51xJCStRWpX2mXOTluIC2CUO_6Umpna3lPL5Es/edit#slide=id.p slides])! This refactoring will allow us to remove a lot of networking logic out of individual model containers.</span> <span lang="en" dir="ltr" class="mw-content-ltr">For example, currently if there was some changing to the discovery.wmnet endpoint (WMF's internal endpoint for APIs), we'd have to update hundreds of individual model containers and redeploy them.</span> <span lang="en" dir="ltr" class="mw-content-ltr">This refactoring removes this need entirely.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We've been deploying AMD's open source software stack (ROCm) inside each k8s node, but we suspect this has been unneccessary (and actually causing some problems) because PyTorch already has a version of ROCm included in the library. This work is being prioritized because completing it is a requiring for making the large GPU order have have planned later in the quarter.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We've been deploying AMD's open source software stack (ROCm) inside each k8s node, but we suspect this has been unnecessary (and actually causing some problems) because PyTorch already has a version of ROCm included in the library. This work is being prioritized because completing it is a requiring for making the large GPU order have have planned later in the quarter.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We are preparing a patch that enables logo-detection model server to access external URLs using internal k8s endpoints. This is part of some of the changes we needed to make to deploy the model.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We are preparing a patch that enables logo-detection model server to access external URLs using internal k8s endpoints. This is part of some of the changes we needed to make to deploy the model.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Continuing to test the HuggingFace model server image on our Lift Wing nodes. This work was paused for the week while the engineer attended the Wikipedia Hackathon in Tallinn.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Continuing to test the HuggingFace model server image on our Lift Wing nodes. This work was paused for the week while the engineer attended the Wikipedia Hackathon in Tallinn.</span>
Line 60: Line 76:
* {{ymd|2024|4|26}}
* {{ymd|2024|4|26}}
** <span lang="en" dir="ltr" class="mw-content-ltr">Reviewing and testing the big patch for the ORES extension. The ORES extension provides a way to see the probability that a particular edit is reverted for all edits on the recent changes page of many Wikis. The new revert risk model into the extension so that volunteers can use that new model when hunting down potential vandalism.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Reviewing and testing the big patch for the ORES extension. The ORES extension provides a way to see the probability that a particular edit is reverted for all edits on the recent changes page of many Wikis. The new revert risk model into the extension so that volunteers can use that new model when hunting down potential vandalism.</span>
* <span lang="en" dir="ltr" class="mw-content-ltr">We're still doing some tweaks for the image processing for the logo detection model, specifically restricting the image processing to trusted domains that host Wikimedia comments images.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We're still doing some tweaks for the image processing for the logo detection model, specifically restricting the image processing to trusted domains that host Wikimedia comments images.</span>
* <span lang="en" dir="ltr" class="mw-content-ltr">We have a big Istio (Istio is the service mesh for k8s that controls how microservices share data with each other) refactoring proposal under discussion. On Tuesday the team will have a special meeting to discuss the proposed refactoring and decide on the path forward. I'll post the slides next week if people are interested.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We have a big Istio (Istio is the service mesh for k8s that controls how microservices share data with each other) refactoring proposal under discussion. On Tuesday the team will have a special meeting to discuss the proposed refactoring and decide on the path forward. I'll post the slides next week if people are interested.</span>
* {{ymd|2024|4|19}}
* {{ymd|2024|4|19}}
** <span lang="en" dir="ltr" class="mw-content-ltr">The logo detection model is being moved to the experimental namespace. This will be a moment where we can test the model in a production setting to make sure that it has the performance that we want. This work is being coordinated really closely with the structured content team to make sure it meets their needs.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">The logo detection model is being moved to the experimental namespace. This will be a moment where we can test the model in a production setting to make sure that it has the performance that we want. This work is being coordinated really closely with the structured content team to make sure it meets their needs.</span>
Line 75: Line 91:
** <span lang="en" dir="ltr" class="mw-content-ltr">Work on revertrisk-multilingual GPU image, ensure the RRML model is compatible with torch 2.x (e.g. predictions are correct as the model was trained with 1.13)</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Work on revertrisk-multilingual GPU image, ensure the RRML model is compatible with torch 2.x (e.g. predictions are correct as the model was trained with 1.13)</span>
* {{ymd|2024|3|29}}
* {{ymd|2024|3|29}}
** <span lang="en" dir="ltr" class="mw-content-ltr">We are still working on the logo detection model for Wikimedia Commons. The current status is that we have confirmed with the product team working on the feature that the model is returning the expected outputs. The next step is to look at input validation and image size limits. The open question we are discussing with the product team is whether resizing of images should be done inside Lift Wing or prior to the image being sent to Lift Wing. Resizing is important because the logo detection model expects an image of a certain size.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We are still working on the logo detection model for Wikimedia Commons.</span> <span lang="en" dir="ltr" class="mw-content-ltr">The current status is that we have confirmed with the product team working on the feature that the model is returning the expected outputs.</span> <span lang="en" dir="ltr" class="mw-content-ltr">The next step is to look at input validation and image size limits.</span> <span lang="en" dir="ltr" class="mw-content-ltr">The open question we are discussing with the product team is whether resizing of images should be done inside Lift Wing or prior to the image being sent to Lift Wing.</span> <span lang="en" dir="ltr" class="mw-content-ltr">Resizing is important because the logo detection model expects an image of a certain size.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Work / banging our heads continues on the pytorch base image. For those following along, we are working with Service Ops to make a reasonably sized docker image that contains pytorch and ROCm support. If the base image is too big it becomes a problem for our Docker registry and we are trying to be good stewards of that common resource. Turns out it is harder than we thought.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Work / banging our heads continues on the pytorch base image.</span> <span lang="en" dir="ltr" class="mw-content-ltr">For those following along, we are working with Service Ops to make a reasonably sized docker image that contains pytorch and ROCm support.</span> <span lang="en" dir="ltr" class="mw-content-ltr">If the base image is too big it becomes a problem for our Docker registry and we are trying to be good stewards of that common resource.</span> <span lang="en" dir="ltr" class="mw-content-ltr">Turns out it is harder than we thought.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">More work is happening on Lift Wing caching. We are still working out how we want Lift Wing (specifically KServe's Istio) to talk to the Cassandra servers.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">More work is happening on Lift Wing caching. We are still working out how we want Lift Wing (specifically KServe's Istio) to talk to the Cassandra servers.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">A new version of the Language Agnostic Revert Risk model has been deployed to staging and is currently doing load testing.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">A new version of the Language Agnostic Revert Risk model has been deployed to staging and is currently doing load testing.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">More work on the HuggingFace model server integration with Lift Wing. Once we crack this we will be able to deploy most models on HuggingFace quickly.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">More work on the HuggingFace model server integration with Lift Wing. Once we crack this we will be able to deploy most models on HuggingFace quickly.</span>
* {{ymd|2024|3|22}}
* {{ymd|2024|3|22}}
** <span lang="en" dir="ltr" class="mw-content-ltr">We stood up a Wikimedia community of practice for ML this week. The goal is to provide a space for all the folks around WMF that are working on the technical side of ML to share insights and learn together. Currently there are folks from a number of teams in the community of practice, including ML, Research, Content Translation, and others.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We stood up a Wikimedia community of practice for ML this week.</span> <span lang="en" dir="ltr" class="mw-content-ltr">The goal is to provide a space for all the folks around WMF that are working on the technical side of ML to share insights and learn together.</span> <span lang="en" dir="ltr" class="mw-content-ltr">Currently there are folks from a number of teams in the community of practice, including ML, Research, Content Translation, and others.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We are still waiting for our test GPUs (one server with two MI210s) to be installed in the data center. Once we test this configuration works well in our infrastructure (a few days of testing max) we can continue with the full order.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We are still waiting for our test GPUs (one server with two MI210s) to be installed in the data center. Once we test this configuration works well in our infrastructure (a few days of testing max) we can continue with the full order.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">I am starting work on a white paper that surveys all the work Wikimedia'verse is doing around AI, this includes models WMF hosts, advocacy work done by WMF, work by volunteers, etc. If you know some people I should talk with, definitely reach out.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">I am starting work on a white paper that surveys all the work Wikimedia'verse is doing around AI, this includes models WMF hosts, advocacy work done by WMF, work by volunteers, etc. If you know some people I should talk with, definitely reach out.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We are really pushing hard on getting caching deployed. The reason is that with caching, it means we can really take full advantage of the CPUs we have now by pre-caching predictions. The end result for users is that a prediction that might take 500ms would take a fraction of that time. The exact current status of the work is that our SRE is trying to get Lift Wing to speak to the Cassandra servers.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We are really pushing hard on getting caching deployed.</span> <span lang="en" dir="ltr" class="mw-content-ltr">The reason is that with caching, it means we can really take full advantage of the CPUs we have now by pre-caching predictions.</span> <span lang="en" dir="ltr" class="mw-content-ltr">The end result for users is that a prediction that might take 500ms would take a fraction of that time.</span> <span lang="en" dir="ltr" class="mw-content-ltr">The exact current status of the work is that our SRE is trying to get Lift Wing to speak to the Cassandra servers.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Our SLO [https://grafana.wikimedia.org/d/slo-Lift_Wing_Revert_Risk_LA/lift-wing-revert-risk-la-slo-s?orgId=1 dashboards] need to be fixed. They are giving some wild numbers that are clearly incorrect. Our team is working with folks to figure it out.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Our SLO [https://grafana.wikimedia.org/d/slo-Lift_Wing_Revert_Risk_LA/lift-wing-revert-risk-la-slo-s?orgId=1 dashboards] need to be fixed. They are giving some wild numbers that are clearly incorrect. Our team is working with folks to figure it out.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Work on the Logo Detection model continues. The request to host this model comes from the Structured Content team. The goal is to [https://phabricator.wikimedia.org/T358676 predict logos in Wikimedia Commons] because logos account for a significant chunk of files that receive a deletion request.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Work on the Logo Detection model continues.</span> <span lang="en" dir="ltr" class="mw-content-ltr">The request to host this model comes from the Structured Content team.</span> <span lang="en" dir="ltr" class="mw-content-ltr">The goal is to [[phab:T358676|predict logos in Wikimedia Commons]] because logos account for a significant chunk of files that receive a deletion request.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We are continuing to try to load the HuggingFace model server onto Lift Wing. When completed this offers the potential to load a model hosted on HuggingFace into Lift Wing quickly and easily, opening a huge new library of models for folks to use.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We are continuing to try to load the HuggingFace model server onto Lift Wing. When completed this offers the potential to load a model hosted on HuggingFace into Lift Wing quickly and easily, opening a huge new library of models for folks to use.</span>
* {{ymd|2024|3|15}}
* {{ymd|2024|3|15}}
** <span lang="en" dir="ltr" class="mw-content-ltr">We are working on deploying a model for the Structured Content team that detects potentially copyrighted image uploads on Commons, specifically images with logos. ([https://phabricator.wikimedia.org/T358676 T358676])</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We are working on deploying a model for the Structured Content team that detects potentially copyrighted image uploads on Commons, specifically images with logos.</span> ([[phab:T358676|T358676]])
** <span lang="en" dir="ltr" class="mw-content-ltr">We are continuing to work on hosting HuggingFace model server on Lift Wing. This would make deploying HuggingFace models super simple.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We are continuing to work on hosting HuggingFace model server on Lift Wing. This would make deploying HuggingFace models super simple.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We have deployed Dragonfly cache on Lift Wing to help with Docker image sizes.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We have deployed Dragonfly cache on Lift Wing to help with Docker image sizes.</span>
Line 98: Line 114:
** <span lang="en" dir="ltr" class="mw-content-ltr">An issue we are facing is that WMF's docker registry is set up for smaller docker images (~2GBs). However, the docker images of the team can get pretty big because of ROCm/Pytorch (~6-8GB). We are working out how to resolve that. There a number of strategies can do, from optimizing the image layers better to requesting the max docker image size limit to be increased.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">An issue we are facing is that WMF's docker registry is set up for smaller docker images (~2GBs). However, the docker images of the team can get pretty big because of ROCm/Pytorch (~6-8GB). We are working out how to resolve that. There a number of strategies can do, from optimizing the image layers better to requesting the max docker image size limit to be increased.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">As a partial solution to the above, we installed Dragonfly, which is a peer-2-peer layer between our Kubernetes cluster and the WMF docker registry. We will also work on some other improvements.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">As a partial solution to the above, we installed Dragonfly, which is a peer-2-peer layer between our Kubernetes cluster and the WMF docker registry. We will also work on some other improvements.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We are continue working on including HuggingFace's prebuilt model server into Lift Wing. This would mean we could quickly deploy any model on HuggingFace with all the optimizations HuggingFace provides. ([https://phabricator.wikimedia.org/T357986 T357986]). This isn't done yet but it would be really nice to have.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We are continue working on including HuggingFace's prebuilt model server into Lift Wing. This would mean we could quickly deploy any model on HuggingFace with all the optimizations HuggingFace provides.</span> ([[phab:T357986|T357986]]) <span lang="en" dir="ltr" class="mw-content-ltr">This isn't done yet but it would be really nice to have.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Fixing a bug reported about inconsistent data type for article quality scores on ptwiki. The error as because of the mixed schema of the responses returned by ORES.([https://phabricator.wikimedia.org/T358953 T358953])</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Fixing a bug reported about inconsistent data type for article quality scores on ptwiki. The error as because of the mixed schema of the responses returned by ORES.</span> ([[phab:T358953|T358953]])
** <span lang="en" dir="ltr" class="mw-content-ltr">We made our server hardware request for the next fiscal year. The short version is: GPUs.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">We made our server hardware request for the next fiscal year. The short version is: GPUs.</span>
* {{ymd|2024|3|1}}
* {{ymd|2024|3|1}}
** <span lang="en" dir="ltr" class="mw-content-ltr">GPU order is underway. We are in the process of ordering a series of servers to use for training and inference. Each server will have two MI210 AMD GPUs. Most will be reserved for model inference (specifically, larger models like LLMs), but we will use two servers (4 GPUs) to create a model training environment. This model training environment will start very small and scrappy but will hopefully grow into a place for automated retraining of models and the standardization of model training approaches. The next steps are a single server will on its way to our data center, once this is tested we will make the full order.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">GPU order is underway.</span> <span lang="en" dir="ltr" class="mw-content-ltr">We are in the process of ordering a series of servers to use for training and inference.</span> <span lang="en" dir="ltr" class="mw-content-ltr">Each server will have two MI210 AMD GPUs.</span> <span lang="en" dir="ltr" class="mw-content-ltr">Most will be reserved for model inference (specifically, larger models like LLMs), but we will use two servers (4 GPUs) to create a model training environment.</span> <span lang="en" dir="ltr" class="mw-content-ltr">This model training environment will start very small and scrappy but will hopefully grow into a place for automated retraining of models and the standardization of model training approaches.</span> <span lang="en" dir="ltr" class="mw-content-ltr">The next steps are a single server will on its way to our data center, once this is tested we will make the full order.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Work on caching for Lift Wing continues. We have in the process of making a large order of GPUs. However, to optimize our resource use, one of the best strategies we can do is conduct model inference using our existing CPUs. This is not always possible, for example cases when the set of possible model inputs is not finite. However, in cases where the possible inputs are finite we can cache the predictions for those inputs and then serve them to users rapidly with minimal compute used. This is a similar system to that which was originally used on ORES.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Work on caching for Lift Wing continues.</span> <span lang="en" dir="ltr" class="mw-content-ltr">We have in the process of making a large order of GPUs.</span> <span lang="en" dir="ltr" class="mw-content-ltr">However, to optimize our resource use, one of the best strategies we can do is conduct model inference using our existing CPUs.</span> <span lang="en" dir="ltr" class="mw-content-ltr">This is not always possible, for example cases when the set of possible model inputs is not finite.</span> <span lang="en" dir="ltr" class="mw-content-ltr">However, in cases where the possible inputs are finite we can cache the predictions for those inputs and then serve them to users rapidly with minimal compute used.</span> <span lang="en" dir="ltr" class="mw-content-ltr">This is a similar system to that which was originally used on ORES.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">The pentesting of Lift Wing continues. The testing is being done by a third party contractor and is examining our vulnerability to malicious code.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">The pentesting of Lift Wing continues. The testing is being done by a third party contractor and is examining our vulnerability to malicious code.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Wikimedia's branding team has come out with some suggestions for the naming of machine learning tools and models. The hope is that our naming is more systematic and less ad-hoc.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Wikimedia's branding team has come out with some suggestions for the naming of machine learning tools and models. The hope is that our naming is more systematic and less ad-hoc.</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Chris helped organize and attend an event in Bellagio, Italy to craft a research agenda for researchers interested in Wikipedia. That research agenda is [[metawiki:Artificial_intelligence/Bellagio_2024|available here]].</span>
** <span lang="en" dir="ltr" class="mw-content-ltr">Chris helped organize and attend an event in Bellagio, Italy to craft a research agenda for researchers interested in Wikipedia. That research agenda is [[m:Artificial intelligence/Bellagio 2024|available here]].</span>

[[Category:Research{{#translation:}}]]

Latest revision as of 11:42, 24 February 2026

Willkommen auf der Homepage des Wikimedia Foundation Teams für Maschinelles Lernen

Unser Team kümmert sich um die Entwicklung und Verwaltung von Machine-Learning-Modellen für Endnutzer sowie um die Infrastruktur, die für die Entwicklung, das Training und den Einsatz dieser Modelle erforderlich ist.

Aktuelle Projekte

For archived projects, see this list.

Kontaktiere uns

Du hast eine Frage? Möchtest du mit dem Team oder unserer Gemeinschaft von Freiwilligen über maschinelles Lernen sprechen? Hier sind die besten Möglichkeiten, mit uns in Kontakt zu treten.

Team Chat

Diskutiere über maschinelles Lernen und beobachte das Team bei der Arbeit in unserem öffentlichen IRC-Chatroom #wikimedia-ml connect auf irc.libera.chat.

Aktive Arbeitspinnwand

Wenn du ein bestimmtes Anliegen hast, das du diskutieren oder bearbeiten möchtest, tritt unserem öffentlichen Board Phabricator bei. Besuche unser Board

Was ist neu?

    • We are continuing to integrate the article-country model into Liftwing. The article country model predicts which countries any particular model will be applicable for, and it's an extension of the article-topic model, which we have used for years.
    • We're trying different approaches to build vllm (a high-throughput and memory-efficient system designed for serving large language models) and ROCm (the code that allows the CPU to talk to AMD GPUs) with Ubuntu. This is part of the work of making production LLMs on Liftwing possible.
    • We're currently working on configuring the ML Lab servers. These are for model training.
    • Updated the rec-api image deployment model. Deployed the reference need model to production.
    • Following up on recurring issue reported by the Structured Content team: The MediaDetection API can access the logo-detection endpoint via mwdebug1001.eqiad.wmnet and mwdebug2001.codfw.wmnet, but can't access it on k8s-mwdebug
    • Adding logo-detection documentation to the API portal docs.
    • Investigating occasional slow queries on LiftWing when using some RevScoring models
    • Continuing the remaining work on the pre-save revertrisk model. This model is designed to provide a vandalism prediction before an edit is saved to Wikipedia (and thus doesn't have a revision ID)
    • Work continues on upgrading kserve to 0.13
    • Initializing install config for the GPU hosts in eqiad
    • Apologies on my end for the delay in updates, I got covid.
    • Working continues on the Logo Detection model. We made an example logo-detection model-server that processes base64 image objects instead of image URLs and sent it to the Structured Content team for their thoughts.
    • Work continues on the HuggingFace model server.
    • General bug fixes and improvements.
    • Die Arbeit am Logo-Erkennungsmodell geht weiter. Diese Woche haben wir im Team darüber diskutiert, ob das kodierte Bild direkt an LiftWing gesendet werden soll oder nicht. Alternativ würden wir eine URL mit dem Speicherort des Bildes erhalten, auf die LiftWing dann zugreifen und es herunterladen kann. Das ist wichtig, weil es sich auf die Größe der REST-Nutzdaten auswirkt, insbesondere bei gebündelten Anfragen.
    • We are still working on the HuggingFace model server GPU issue (i.e. it won't recognize our AMD GPU). There are a number of possibilities as to why, but we want this resolved before we finalize our order for this fiscal year.
    • A number of misc bug fixes and improvements.
    • Our big Istio refactoring is underway (slides)! This refactoring will allow us to remove a lot of networking logic out of individual model containers. For example, currently if there was some changing to the discovery.wmnet endpoint (WMF's internal endpoint for APIs), we'd have to update hundreds of individual model containers and redeploy them. This refactoring removes this need entirely.
    • We've been deploying AMD's open source software stack (ROCm) inside each k8s node, but we suspect this has been unnecessary (and actually causing some problems) because PyTorch already has a version of ROCm included in the library. This work is being prioritized because completing it is a requiring for making the large GPU order have have planned later in the quarter.
    • We are preparing a patch that enables logo-detection model server to access external URLs using internal k8s endpoints. This is part of some of the changes we needed to make to deploy the model.
    • Continuing to test the HuggingFace model server image on our Lift Wing nodes. This work was paused for the week while the engineer attended the Wikipedia Hackathon in Tallinn.
    • Lift Wing caching work has been paused until the Istio refactoring is complete.
    • Reviewing and testing the big patch for the ORES extension. The ORES extension provides a way to see the probability that a particular edit is reverted for all edits on the recent changes page of many Wikis. The new revert risk model into the extension so that volunteers can use that new model when hunting down potential vandalism.
    • We're still doing some tweaks for the image processing for the logo detection model, specifically restricting the image processing to trusted domains that host Wikimedia comments images.
    • We have a big Istio (Istio is the service mesh for k8s that controls how microservices share data with each other) refactoring proposal under discussion. On Tuesday the team will have a special meeting to discuss the proposed refactoring and decide on the path forward. I'll post the slides next week if people are interested.
    • The logo detection model is being moved to the experimental namespace. This will be a moment where we can test the model in a production setting to make sure that it has the performance that we want. This work is being coordinated really closely with the structured content team to make sure it meets their needs.
    • ML and research Airflow Pipeline Sprint has started this week. This is a effort to see how we can use Airflow pipelines and GPS on the existing Hadoop Cluster to train models.
    • Work continues on the Cassandra clusters that will be part of the caching solution.
    • Work continues on the Hugging Face model server image. This is an effort that we're working on that will allow us to easily host many of the models that are available on Hugging Face onto Lift Wing directly. This is actually a really interesting project because it's an easy way for the community to experiment with the models that they might want to host on Lift Wing and even propose models that they might want to have on Lift Wing.
    • We are working with the data center operations team on the procurement of new machines with GPUs. The current status is that we are working with the vendor to an issue around the availability of a particular server configuration and looking at some alternatives.
    • Chris on vacation. No update this week.
    • Big win for the week: Our HuggingFace Docker image patch has been reviewed and approved. This Docker image allows us to deploy HuggingFace models quickly onto LiftWing, in a way that will speed up all development process going forward.
    • Continuing to integrate the logo-detection prototype into KServe custom model-server that will be hosted on LiftWing
    • Work on revertrisk-multilingual GPU image, ensure the RRML model is compatible with torch 2.x (e.g. predictions are correct as the model was trained with 1.13)
    • We are still working on the logo detection model for Wikimedia Commons. The current status is that we have confirmed with the product team working on the feature that the model is returning the expected outputs. The next step is to look at input validation and image size limits. The open question we are discussing with the product team is whether resizing of images should be done inside Lift Wing or prior to the image being sent to Lift Wing. Resizing is important because the logo detection model expects an image of a certain size.
    • Work / banging our heads continues on the pytorch base image. For those following along, we are working with Service Ops to make a reasonably sized docker image that contains pytorch and ROCm support. If the base image is too big it becomes a problem for our Docker registry and we are trying to be good stewards of that common resource. Turns out it is harder than we thought.
    • More work is happening on Lift Wing caching. We are still working out how we want Lift Wing (specifically KServe's Istio) to talk to the Cassandra servers.
    • A new version of the Language Agnostic Revert Risk model has been deployed to staging and is currently doing load testing.
    • More work on the HuggingFace model server integration with Lift Wing. Once we crack this we will be able to deploy most models on HuggingFace quickly.
    • We stood up a Wikimedia community of practice for ML this week. The goal is to provide a space for all the folks around WMF that are working on the technical side of ML to share insights and learn together. Currently there are folks from a number of teams in the community of practice, including ML, Research, Content Translation, and others.
    • We are still waiting for our test GPUs (one server with two MI210s) to be installed in the data center. Once we test this configuration works well in our infrastructure (a few days of testing max) we can continue with the full order.
    • I am starting work on a white paper that surveys all the work Wikimedia'verse is doing around AI, this includes models WMF hosts, advocacy work done by WMF, work by volunteers, etc. If you know some people I should talk with, definitely reach out.
    • We are really pushing hard on getting caching deployed. The reason is that with caching, it means we can really take full advantage of the CPUs we have now by pre-caching predictions. The end result for users is that a prediction that might take 500ms would take a fraction of that time. The exact current status of the work is that our SRE is trying to get Lift Wing to speak to the Cassandra servers.
    • Our SLO dashboards need to be fixed. They are giving some wild numbers that are clearly incorrect. Our team is working with folks to figure it out.
    • Work on the Logo Detection model continues. The request to host this model comes from the Structured Content team. The goal is to predict logos in Wikimedia Commons because logos account for a significant chunk of files that receive a deletion request.
    • We are continuing to try to load the HuggingFace model server onto Lift Wing. When completed this offers the potential to load a model hosted on HuggingFace into Lift Wing quickly and easily, opening a huge new library of models for folks to use.
    • We are working on deploying a model for the Structured Content team that detects potentially copyrighted image uploads on Commons, specifically images with logos. (T358676)
    • We are continuing to work on hosting HuggingFace model server on Lift Wing. This would make deploying HuggingFace models super simple.
    • We have deployed Dragonfly cache on Lift Wing to help with Docker image sizes.
    • Our Cassandra databases for an eventual caching system is in production. Still more work to do but its a good start.
    • General updates and bug fixes.
    • Sorry for the update being one day late, Chris (I) attended the Strategy meeting in NYC and is writing this update from the plane back.
    • An issue we are facing is that WMF's docker registry is set up for smaller docker images (~2GBs). However, the docker images of the team can get pretty big because of ROCm/Pytorch (~6-8GB). We are working out how to resolve that. There a number of strategies can do, from optimizing the image layers better to requesting the max docker image size limit to be increased.
    • As a partial solution to the above, we installed Dragonfly, which is a peer-2-peer layer between our Kubernetes cluster and the WMF docker registry. We will also work on some other improvements.
    • We are continue working on including HuggingFace's prebuilt model server into Lift Wing. This would mean we could quickly deploy any model on HuggingFace with all the optimizations HuggingFace provides. (T357986) This isn't done yet but it would be really nice to have.
    • Fixing a bug reported about inconsistent data type for article quality scores on ptwiki. The error as because of the mixed schema of the responses returned by ORES. (T358953)
    • We made our server hardware request for the next fiscal year. The short version is: GPUs.
    • GPU order is underway. We are in the process of ordering a series of servers to use for training and inference. Each server will have two MI210 AMD GPUs. Most will be reserved for model inference (specifically, larger models like LLMs), but we will use two servers (4 GPUs) to create a model training environment. This model training environment will start very small and scrappy but will hopefully grow into a place for automated retraining of models and the standardization of model training approaches. The next steps are a single server will on its way to our data center, once this is tested we will make the full order.
    • Work on caching for Lift Wing continues. We have in the process of making a large order of GPUs. However, to optimize our resource use, one of the best strategies we can do is conduct model inference using our existing CPUs. This is not always possible, for example cases when the set of possible model inputs is not finite. However, in cases where the possible inputs are finite we can cache the predictions for those inputs and then serve them to users rapidly with minimal compute used. This is a similar system to that which was originally used on ORES.
    • The pentesting of Lift Wing continues. The testing is being done by a third party contractor and is examining our vulnerability to malicious code.
    • Wikimedia's branding team has come out with some suggestions for the naming of machine learning tools and models. The hope is that our naming is more systematic and less ad-hoc.
    • Chris helped organize and attend an event in Bellagio, Italy to craft a research agenda for researchers interested in Wikipedia. That research agenda is available here.