Wikisource:Scriptorium: Difference between revisions
archiving |
|||
| Line 127: | Line 127: | ||
::Yeah, I think Brion did an import of a number of Sanskrit pages to test out the function. But to this moment, the import feature still only allows us to import pages to the Sanskrit Wikisource, nowhere else.{{User:Zhaladshar/sig}} 19:53, 15 August 2005 (UTC) |
::Yeah, I think Brion did an import of a number of Sanskrit pages to test out the function. But to this moment, the import feature still only allows us to import pages to the Sanskrit Wikisource, nowhere else.{{User:Zhaladshar/sig}} 19:53, 15 August 2005 (UTC) |
||
:::Sorry, I'd just assumed this was going ahead now and I'm not sure what the hold up is. Did the Sanskrit Wikisource ever exist? Neither Tim or Brion are on IRC at the moment, but I'll try to find out more when they are. [[User:Angela|Angela]] 18:54, 16 August 2005 (UTC) |
|||
Revision as of 18:54, 16 August 2005
Edit(+) / Seite bearbeiten(+) / Modifier(+)
(= Wikipedia:Village pump - Wikipedia:Forum - - bar di Wikipedia)
Wikisource project pages
- What is Wikisource?
- Wikisource and Wikibooks
- Votes
- Statistics
- Copyright
- Possible copyright violations
- Proposed deletions
- Cataloging
- Wikisource and Project Gutenberg
- Languages - see below
Separate discussions
Verschobene Diskussionen / Discussions déplacées / verschoven besprekingen / 移动的讨论 / Discussioni spostate
Logo
Languages
- Language policy
- Language domain requests
- Language domain requests/Rules for voting
- List of Wikisource Languages (=Main Page)
The language domain vote (over now):
- New vote on language subdomains (vote has concluded)
- New discussion on language subdomains (vote has concluded)
Archives
Archive / 存 / Archivi (oldest first)
- m:Wikisource (Pre-launch discussions)
- November 2003 to January 2004
- Feb 2004 - Jun 2004
- Jul 2004 - Oct 2004
- Oct 2004 - Feb 2005
- Feb 2005 - Apr 2005
- Apr 2005 - Aug 2005
- ==================================================
Wikibooks Card Catalog Office
I've started a little project (if you can call it that) over at Wikibooks and I'd like some input from Wikisource as well on a certain aspect of my proposal as well. The link to the project is: b:Wikibooks:Card Catalog Office
In going over some of the common mistakes for new users on Wikibooks, we've noticed a large number of raw source texts that often get simply dumped onto Wikibooks. In a bid to help point out that Wikisource does exist, we are suggesting that we integrate both Wikibooks and Wikisource into this cataloging system... and perhaps even make it Wikimedia-wide in general for similar content. Wikibooks and Wikisource are the ones most affected by this proposal.
Essentially, you can have a "one-stop shopping experience" to find a "book" in the "library" to all Wikimedia content. Of course (I don't know if this is getting snobbish or not) all of Wikipedia would be only one of the "entries" in this library catalog, as would Wiktionary. This is book-style content that would be cataloged, and it would include cataloging systems like the Library of Congress and Dewey Decimal type cataloging, in addition to custom cataloging systems developed just for Wikimedia. The category systems for each project (Wikibooks and Wikisource) would stay seperated, and I'm not asking for content to be moved either. This is just a linking guide to combine an index between the two projects.
I'm also willing to accept if the "regulars" here on Wikisource think this is a bad idea, but I would like to get some input. There are some technical aspects that certainly need to be coordinated or even invented, but I think it could benefit both projects and draw help and assistance from a wider audience as well. --Robert Horning 19:06, 5 August 2005 (UTC)
- It's an interesting proposal. I assume then that if you looked up "Chocolate" you would find: articles relating to chocolate, recipes relating to chocolate, source texts about chocolate, images of chocolate, news about chocolate etc. A "cross-wiki" way of categorising as such? You would probably run into problems with systems such as the Dewey Decimal system due to the fact that new media like Computer Games would not be adequately covered. I do like the idea (in the same way you can access the different searches on Google), I just think it might need a bit more fleshing out before I'm fully convinced ;) Greg Robson 09:25, 6 August 2005 (UTC)
- I don't think any one particular system of organization or searching is adequate in every situation. Google, for instance, is a good word search system when you know what you are trying to find. It really doesn't do a good job of showing slightly related web pages that may not use the exact wording you used for the search term. Library cataloging systems can do that to some degree (they aren't perfect either) and there are other approaches including the catagories within each Wikimedia project.
Right now we are trying to get things organized, so if there is anybody here that would like to help out, they are invited as well. The long-term goal is to get Wikisource included, as Wikisource does have book-like materials that are similar in nature to Wikibooks, and can also be accomdated by a structure like this. I'll give more feedback later once we got a signficant portion of Wikibooks organized. What is interesting is that this isn't really the first effort to do this, but there have been some systemic problems with the other approaches in the past that we are trying to overcome, including problems of trying to maintain a huge number of pages that generally were considered more of a chore to take care of than something anybody was interested in working on. --Robert Horning 11:52, 6 August 2005 (UTC)
- I don't think any one particular system of organization or searching is adequate in every situation. Google, for instance, is a good word search system when you know what you are trying to find. It really doesn't do a good job of showing slightly related web pages that may not use the exact wording you used for the search term. Library cataloging systems can do that to some degree (they aren't perfect either) and there are other approaches including the catagories within each Wikimedia project.
- I'm a bit unsure of how I can be of help. Usually I don't work on other projects, as I've got my hands full here, but as this is something that will affect Wikisource greatly, I'm willing to work on the project. | 12:32, 6 August 2005 (UTC)
Oral Recordings of Texts at Wikisource
Does anyone have any ideas on how Wikisource might incorporate oral recordings of the texts that are kept here? My initial thought is that while such recordings should obviously be uploaded at Commons (not here), the texts here can and should have direct links to them. By clicking a link, one could listen to an oral reading while viewing the text on the screen.Dovi 12:27, 7 August 2005 (UTC)
- I think you know w:en:Wikipedia:WikiProject Spoken Wikipedia? People at English Wikipedia seem to place the recordings at their Wiki, but I think the Commons should be the right place for it. --Jofi 21:56:10, 2005-08-08 (UTC)
- I can understand the recording of a Wikipedia article being placed on en.wikipedia, because it is only useful for that wiki alone. I was thinking along different lines: Suppose we have a Latin text at the Latin Wikisource, a Latin-French bilingual edition at the French Wikisource, and a Latin-English edition at the English Wikisource. A recording of the Latin text would have to kept at the Commons, so that it can be linked to from all three Wikisources (this consideration would be irrelevant to Wikipedia articles, because the language versions are not identical or even close).
- What I am trying to think through is this: Just like all the Wikipedias classify Commons' media for encyclopedic uses, so to all the Wikisources could/should classify source-text-recordings on Commons for library purposes. Our goal, after all, is to be a library of published texts, and good libraries also hold recordings.Dovi 19:14, 10 August 2005 (UTC)
- I don't think it would be useful if the same (supposed) Latin text were at Latin Wikisource and at English, French, ... Wikisource together with the translation. I would prefer a software solution, where it is possible to select two language versions of a text and display them together in two seperate columns. I already thought about making this a feature request but since the developers obviously don't have the time to set up the Wikisources, I didn't discussed the idea yet.
- But besides this it will be possible to link to any Wikimedia file from every project, so it doesn't really matter if the recordings are at the Commons or the appropriate Wikisource. --Jofi 20:30:42, 2005-08-11 (UTC)
- There is no conflict with the two-column software solution. Everyone would love a feature like that, and it could be used as above, only better! What I was talking about was an additional feature, that for a bilingual text an oral reading could be linked to for both languages.Dovi 17:18, 15 August 2005 (UTC)
PDFs acceptable?
I have a crapload of public domain PDFs of interesting articles, government documents, and things like that. Are PDFs acceptable here? I don't have the time or inclination to type out the contents personally, and some of them are many dozens of pages long. They don't seem quite appropriate for Commons, which seems to include all media except text. --Fastfission 00:14, 9 August 2005 (UTC)
- No, PDFs as such should not be uploaded here. Sometimes you can extract the content with pdftotext (on Linux) or equivalent tools and post the content here. Otherwise OCR should be used to get the text. Yann 07:53, 9 August 2005 (UTC)
- As well, if you use Adobe, you can hover the cursor over the text and select it--just like you would do to select text on a web page. It cuts down drastically on the time it takes to transfer the contents of the PDF to a website. | 14:21, 9 August 2005 (UTC)
- Although I don't know Adobe, I think it is only possible if the PDF is made from text, but but not if it is made of images, like [1] ? Yann 19:06, 11 August 2005 (UTC)
- Right. I guess I should have clarified. If the PDF is an image (even an image of text), you cannot select the text. You can only select the text from the PDF if it was a text document or something of the sort made into a PDF. | 19:11, 11 August 2005 (UTC)
- As well, if you use Adobe, you can hover the cursor over the text and select it--just like you would do to select text on a web page. It cuts down drastically on the time it takes to transfer the contents of the PDF to a website. | 14:21, 9 August 2005 (UTC)
- Usually all old texts in PDF files are images and not "real" texts. If the quality of a text doesn't keep perfect if you magnify it, you can be sure it's an image. You can also try the "text selection" button or something like that. If the PDF files are images, the only way to get the text is typing them manually or (usually the better way) using OCR recognition. But if someone has a text as image, that is not freely available at the internet, I think he should be allowed or even be encouraged to place the files here, to give others the possibility to extract the text. --Jofi 20:37:30, 2005-08-11 (UTC)
- That's an idea. We could set up a page listing PDFs that need to be typed and made into real pages here. | 21:00, 11 August 2005 (UTC)
- Copyright could be a problem in some cases: If a text itself is PD but the PDF is from an annotated version, then it wouldn't be possible. But these are most likely rare cases. As far as I know there is no lack of disk space, but if after a period of time nobody did type the PDF, it still could be deleted. --Jofi 21:31:03, 2005-08-11 (UTC)
If the PDF is publicly available somewhere, there is no need to copy it here. But a page with a list of URLs could be useful. Yann 17:57, 12 August 2005 (UTC)
- I've got a number of PDFs that could be copied here--both URLs and files. What if we set up a page like Wikisource:Files to be copied and posted the links there?| 18:04, 12 August 2005 (UTC)
- I think it should rather be PDF/documents that need to be OCRed. Yann 18:39, 12 August 2005 (UTC)
- The PDFs I have, though, don't need to be OCR'd. They're text documents. I know of a million others that are images, however. I guess another name is in order.| 18:53, 12 August 2005 (UTC)
- I think it should rather be PDF/documents that need to be OCRed. Yann 18:39, 12 August 2005 (UTC)
Many of the PDFs are not publicly available elsewhere online. Turning them into straight text would be somewhat difficult — they are often formatted quite poorly (they are old Manhattan Project documents, most of them). Because such sources are only useful if they are reliable, copies of the original images would have to also be able to be viewed as well... so anyway, they don't sound appropriate here, which is too bad. --Fastfission 19:22, 12 August 2005 (UTC)
- What PDFs are you talking about? The ones I have are quite reliable and would be able to be put here.| 19:46, 12 August 2005 (UTC)
- If the PDFs are not available elsewhere, then we can upload them here until they get OCRed and the texts added here. Yann 20:38, 12 August 2005 (UTC)
- Maybe the original PDFs should stay even after they have been added to Wikisource as plain text. So it would be possible to see if there are no mistakes in the OCRed or manually typed texts. If there are problems with disc space some day, the PDFs can still be deleted, but if they are deleted and the user who scanned them is not here anymore, it wouldn't be possible to see if there are mistakes in the text. --Jofi 21:47:55, 2005-08-12 (UTC)
Judicial Opinions?
I am pretty sure Judicial Opinions or court opinions (at least in the United States) are public domain. However, Wexis bills millions annually for access to published opinions. How could we make a start getting opinions on wikisource?
- Yes, as there were published elsewhere, they can be added here. Yann 18:36, 12 August 2005 (UTC)
You can try Oyez.com. They have court opinions for free. I'm sure a ".gov" site also has them for free.| 18:46, 12 August 2005 (UTC)
Language Updates
Today it's 3 months since the voting for language subdomains closed. Perhaps at least the languages that want to start from an empty wiki could be already started. --Jofi 21:49:43, 2005-08-12 (UTC)
- I think as soon is ThomasV returns from his holiday this should be done. As far as I understand, everything is ready: Pages can be transferred and we have a list of languages that are ready.Dovi 03:49, 15 August 2005 (UTC)
- At last! I was already wondering what's going on, as there were no notifications for normal and lower users about the latest work. --Schandolf 16:05, 15 August 2005 (UTC)
- First of all, information is available to all (that's wiki)! Secondly, there is no more info (as far as I know) besides the discussions that have been going on here. Thirdly, I agree with you that we really should have at least semi-regular updates about what's going on. I'm going to ask some people.Dovi 17:26, 15 August 2005 (UTC)
- Dovi's right. I haven't heard about what's going lately, either. I really want the English sub-domain set up, but the import function isn't working yet. ThomasV usually handles asking what's going on (I don't understand IRC), but he's on vacation now. Where would we ask those questions?| 18:51, 15 August 2005 (UTC)
- Angela is always extraordinarily helpful and supportive, so I dropped her a note already. Thomas had been in contact with Tim Starling, who may possibly be the developer who sets this up, so it would be good to ask him too for an update. I also can't figure out IRC (though Thomas tried to explain to me how it's quite easy... :-).Dovi 19:33, 15 August 2005 (UTC)
- PS Wasn't there something earlier about someone doing a mass transfer of 750 pages successfully?Dovi 19:40, 15 August 2005 (UTC)
- Yeah, I think Brion did an import of a number of Sanskrit pages to test out the function. But to this moment, the import feature still only allows us to import pages to the Sanskrit Wikisource, nowhere else.| 19:53, 15 August 2005 (UTC)
- Sorry, I'd just assumed this was going ahead now and I'm not sure what the hold up is. Did the Sanskrit Wikisource ever exist? Neither Tim or Brion are on IRC at the moment, but I'll try to find out more when they are. Angela 18:54, 16 August 2005 (UTC)