[01:46:23] Uhhh well I didn't get it yet but let's say I do have it [01:46:42] Oh thanks [14:39:07] <90gq29> is it possible to request a page like .robots.txt (e.g. none of the wiki ui and just a plain text file)? [14:39:47] <90gq29> im trying to add an [.llms.txt ](https://llmstxt.org/) to the wiki [14:47:30] MediaWiki:Robots.txt I think [14:50:26] <90gq29, replying to thewwrnerdguy> i meant adding a page like robots.txt that displays without loading any of the wiki css and stuff [14:50:39] <90gq29> and mediawiki:robots.txt leads nowhere [14:50:45] On the tech level this is easy because it's not much different from robots.txt. [14:51:26] <90gq29, replying to posix_memalign> thats good to hear [14:51:33] <90gq29> are u able to add it or shall i open a phorge task [14:54:02] like ? [14:54:15] <90gq29, replying to thewwrnerdguy> ye [14:54:23] You could open a phorge task, but I think this needs to be cleared with the community. I would defer to stewards to decide how this gets implemented. [14:54:39] E.g. it may need an RfC. [14:55:29] <90gq29, replying to posix_memalign> is there a reason it would need an rfc? [14:55:38] <90gq29> id only want it added on the wiki i run [14:56:06] <90gq29> i dont really see how it would affect others unless you mean that it might lead to a lot of steward/tech requests [14:56:39] Her's how I would implement it: if MediaWiki:Llms.txt exists, present its content. If it doesn't, return 404. [14:57:11] One needs to be very careful when dealing with LLMs especially since many in the community want nothing to do with it. [14:58:08] E.g. for the first path, a wiki administrator can tell LLMs that it is okay to crawl cotent on the wiki despite not having control over the copyright of the wiki. A random contributor releasing their content under CC BY NC SA may not want their content to be used by LLMs. [14:58:58] <90gq29, replying to posix_memalign> is it not possible to just configure it manually in this case, rather than adding a whole new system for it [14:59:00] For the second path, as long as the 404 message is sufficiently barren I think it'll be fine. [15:01:04] <90gq29, replying to posix_memalign> also, on this note, does google & other chatbots not already crawl miraheze without problem? meaning llms.txt could be used to explicitly state that they shouldnt crawl it? [15:01:37] There's already https://issue-tracker.miraheze.org/T15476 [15:02:07] Google can crawl Miraheze but be block all the other frontier labs. [15:02:29] Mostly because they are a drain on server resources. [15:03:29] If you are referring to per-wiki configurations instead of automatically serving MediaWiki:robots.txt then I'm inclined to avoid this path because it's a lot of work creating and merging pull requests whenever someone wants to update their thing. [15:05:53] <90gq29, replying to posix_memalign> i see so you think i should just write up a short rfc suggesting "if MediaWiki:Llms.txt exists, present its content. If it doesn't, return 404." [15:06:38] <90gq29> anyone can suggest them right? [15:07:39] I'd wait for stewards to determine the appropriate level of community clearance for such a feature. [15:08:03] <90gq29> like wait for stewards to respond with their opinion here? [15:11:39] Yeah. TBH you can probably put it in your AGENTS.md or leave it in a separate local file and instruct the LLM to fetch it in your AGENTS.md. I don't think llms.txt is widely used at all. At least when I tested it the LLM went straight to the main page. [15:11:56] It's a proposed feature that no one follows it seems. [15:15:02] <90gq29, replying to posix_memalign> yeah but i figure it makes it easier to setup bots for others, and it cant hurt right? [15:15:33] [1/4] So: [15:15:33] [2/4] 1. Supporting this feature can be very controversial. [15:15:34] [3/4] 2. It's not very useful because LLMs don't use it. You may have better luck putting it in a hidden div on the main page. [15:15:34] [4/4] Considering these two things I think we should not put too much time in it. [15:18:33] <90gq29, replying to posix_memalign> it can be useful though, no? plus entirely optional so how would it be controversial? [15:19:11] [1/2] Like I tried 3 different models on 2 different coding agents and all of them just go straight to the main page plus `Special:Version` and sometimes api.php. None of them even tries llms.txt or ai.txt. [15:19:11] [2/2] https://cdn.discordapp.com/attachments/1006789349498699827/1531682438617628843/image.png?ex=6a6a19ee&is=6a68c86e&hm=fe243c543990e554a49366cdf3df979362d894af8333307a1a2e76ad5bea2a56& [15:19:49] <90gq29, replying to posix_memalign> u sure they havent noted down earlier that the wiki/miraheze generally doesnt use llms.txt? [15:20:03] I have not. This is the whole conversation. [15:20:07] Have you tried yourself? [15:20:15] https://discord.com/channels/407504499280707585/1006789349498699827/1531677144382574672 [15:21:22] <90gq29, replying to posix_memalign> i havent, but adding "check llms.txt" is still easier than writing up a whole agents.md [15:21:41] <90gq29> and it could be an adoption problem maybe? few websites use/support it so agents dont bother checking for it [15:22:06] <90gq29> if more websites start using it, it would become a more widespread standard and agents would check it more often [15:25:24] Once there is sufficient adoption we can consider adding it. It is not our job to make llms.txt or ai.txt popular. We should spend our limited time in helping out our wiki communities, and I don't see how this is a good use of our time. If there are enough users wanting it or other members of tech think this is a good idea I'm fine with exploring its feasibility, though. [15:27:03] There are 191 requests to llms.txt in the past 24 hours across Miraheze and 15 requests to ai.txt. In comparison we got 356k robots.txt requests. [15:27:04] <90gq29, replying to posix_memalign> sure but didnt you say earlier that this would be really easy to add on a technical level? https://canary.discord.com/channels/407504499280707585/1006789349498699827/1531675286016364544 [15:27:17] Yes. But not at the community level. [15:30:24] If stewards ever pick this up it'll probably end up being another long-winded discussion just like every single AI discussion we had. At its current state I don't think llms.txt is worth the time. There was previously a discussion about no longer blocking LLMs for certain wikis that opt in to it. Various ideas were thrown around and no definitive conclusion. [15:31:01] <90gq29, replying to posix_memalign> that'd be nice too, what was the outcome of that discussion? [15:31:35] <90gq29> in the end isnt miraheze about giving individual wikis the choice, rather than controlling everything? [15:37:51] No outcomes unfortunately. [15:38:14] We do it where feasible but there are certain things we don't allow, like the Widgets extension. [15:40:16] <90gq29, replying to posix_memalign> widgets extension? [15:40:53] There's no substantive evidence that any LLMs actually read or respect that file so I don't see the point personally? [15:41:16] Yes, it's a very old extension and one that people really shouldn't be using anymore. [15:41:45] It basically allowed you to transclude raw HTML. [15:41:52] (Google specifically does not use LLMs.tx and both OpenAi and Anthrophic have stated that they use robots.txt instead) [15:49:56] well the raw html wasn't the problem, the problem was that it had an rce vulnerability so you could literally do whatever you wanted with Miraheze servers [15:50:27] Yeah, and that's the reason why we don't allow it anymore. [15:50:33] Because, I think it had like two more RCE's after the initial one? [15:54:43] Widget RCEs can come either from the extension itself or Smarty, which also has a few RCE vulnerabilities of its own. [15:55:46] I was talking to Redmin about it a while back. If mustache.php is deemed too risky for the Mustache extension it's possible to switch to Lustache and evaluate everything inside Luasandbox. Missing features can still be added back as well. [16:05:00] [1/2] I don't see a problem with this being done unilaterally where there is not a tangible community to assent, as with many other decisions that tend to be 'just permitted' because of this. Organizational wikis may have cause to do this unilaterally as well depending how they are structured. In a standard community wiki where there is a community that should c [16:05:00] [2/2] onsent to foundational changes, this is where I would seek local discussion. This is a fired from the hip notion favoring to offer the choice unless there is a strong argument why the option should be more regulated, which is where I would pass the matter to a global RfC