Hacker Newsnew | past | comments | ask | show | jobs | submit | hashar's commentslogin

France government has announced NixOS based systems some months ago:

- https://github.com/cloud-gouv/securix a hardened/secured os

- https://github.com/cloud-gouv/bureautix-example an example to use securix to build an office deployment (with some packages https://github.com/cloud-gouv/bureautix-example/blob/main/co... )


I’m looking with so much admiration what France has been building in this space, especially the latest gendarmerie transition into open source OS from windows.

The problem though is that there’s always a strong push by French companies and organizations to reject English at any cost, despite the fact that is the most popular language in Europe.

Yes, I get that this is a way to preserve the French language and they have all the rights to do that (in the end, they’re the ones putting the effort in these initiatives) but imagine how faster these ideas could be shared and diffuse in Europe if French understood this. If all the countries in Europe just keep pushing their nationalist ideology instead of truly believing into a united and solid Europe Union, we won’t go any where.


Hi (Sécurix developer), we wrote our documentation in English and in French: https://cloud-gouv.github.io/securix.

Maintaining both languages is hard but we will do our best.


Thanks for the amazing work, let’s keep pushing for independent EU-based technology!

Thank you for your service! I hope you can cooperate also with other countries, I am specifically thinking of Canada here, given they're dual-language English-French. I'm convinced NixOS is the right way forward.

That's quite a generous decision. Thanks a lot for doing your work in public and in multiple languages.

J’imagine que c’est de plus en plus facile de le faire avec l’aide de l’IA!

> The problem though is that there’s always a strong push by French companies and organizations to reject English at any cost

Interestingly, a French-speaking Canadian who has spent some time in France told me that his right to use the French language is better protected in Canada than it is in France. He thinks that this is a consequence of the reciprocality implied by having more than one official language: French speakers in Canada get the same rights as English speakers, so people are happy to grant rights in return for getting them, whereas in France everyone gets screwed and you end up with a situation in which a French person wanting to do a PhD on French history at a French university has to write their grant application in English.


Quebec gets weird with their french rules. Parents are not allowed to send their kids to English speaking schools if both of the parents did their own schooling in French.

[flagged]


That's extremely weird. I don't know of another place that does that aside from Quebec, and it's extremely illogical on its own merits.

"Extremely illogical". Okay buddy.

What other place do you know of whose context is the same as Quebec?


Belgium, Switzerland, Spain, come easily to my mind. Switzerland is a very interesting counter-example because although the French-speaking cantons (Romandie) have French as the "prevalent language", meaning that public services are almost exclusively in French, in practice most people speak German too because they're pragmatic about language: the German kantons are the center of economic activity and not speaking German is a huge disadvantage. Quebec, par contre, is full of people that wistfully (I had intended to write "wishfully" but the auto complete was fantastic this time) think they can manage with speaking only French and ignore the rest of the world.

Your comment is dripping with racism.

Did you know there are tensions in Switzerland because there is a will to stop kids from learning French while they attend primary school? https://www.edi.admin.ch/fr/newnsb/CbsqmmLpkhDjQgEILMp_P

Did you know Switzerland is located in Europe and shares a border with France? Did you know people in France speak French?

Do you know where Quebec is located?

I’ve seen your kind before.


> Your comment is dripping with racism.

Your paranoia is typical of Quebec. You need professional help, Mon ami.

> Did you know there are tensions in Switzerland because there is a will to stop kids from learning French while they attend primary school?

I lived for a while in Zurich as well, and it was well known that kids are exiting high school being unable to speak the second language they supposedly studied.

> Do you know where Quebec is located?

Yeah, I'm in Montreal right now and have been living there for quite a while.


> Yeah, I'm in Montreal right now and have been living there for quite a while.

shocked Pikachu


No matter what I replied, I suspect it would have confirmed your prejudices.

Belgium, somewhat? :)

Not really. They are located in Europe. Share a border with France. Still have a lot of laws to protect French.

It's people who speak french not being able to make a choice about their children's education, yes it's weird, and no it's not judgemental to wonder why parents can't make a very basic decision in their child's education.

> the latest gendarmerie transition into open source OS from windows

That was 18 years ago: https://fr.wikipedia.org/wiki/GendBuntu


> Yes, I get that this is a way to preserve the French language and they have all the rights to do that (in the end, they’re the ones putting the effort in these initiatives) but imagine how faster these ideas could be shared and diffuse in Europe if French understood this.

Isnt this just baked into the idea? The French are building an OS for internal use for their own interests.

If we then try to extend that or the Dutch OS to other countries we have the original problem of dependence.


I guess we could do with more weird third-path stuff with the Linux world, even if the reason for said weird stuff existing is the French having (le) NIH syndrome.

It's also great when an actual commercial organization with requirements gets into the Linux game, as they don't really yield to 'we could make that printer work but it would be against everything we stand for' style arguments that have cemented problems in Linux userland for decades.


‘United and solid EU’ doesn’t have to involve forcing everyone to use a someone’s native language. Ideally we’d have a common language that’s neutral and used for international communication only. In fact, we already have a working solution; Esperanto has existed for well over 100 years at this point, and continues to be a healthy, living language. People just don’t seem to care that much. It always seems to be either practicality or national heritage, never both at the same time. Yet the common ground has been there all along.

Frenchman here. French companies and organization, in my experience actually really like the English language (at OVH for example, most official docs are in English), but I do think it makes sense for a government funded by people who speak mostly French (ie the French taxpayers) to do stuff in French, the language most of its citizens will be most familiar with.

Putting stuff in English will tend to exclude the less educated, or people who struggle to learn foreign languages.

EDIT: and actually, from looking at the README of https://github.com/cloud-gouv/securix

> Ce README est en français mais le reste du code, les issuers et les PR sont en anglais.

The README is in French but the rest of the code, the issues (they wrote issuers, but that is a typo I suspect) and the PRs are in English.

So. :P


Italianman here, I totally understand and respect your point. I just wish this was a collectively funded initiative, English-first but open to all translations. I’d be so happy if my country understood this and invested in this as well, so kudos to France for leading the tech here!

Je t’assure que ce besoin qu’on certains français de montrer leur capacité de parler l’anglais, de nommer leurs entreprises avec des noms à connotation anglophone, etc. est considéré très gênant ici du Québec.

Oui, c'est bien connu que vous Québécois préférez la stratégie de l'autruche.

Tes interventions montrent que tu es un raciste intolérant. Je n’ai plus de temps à perdre avec toi.

Quel dommage.

The Dutch have ditched even github. They are thinking ahead.

https://code.overheid.nl/MinBZK/DAWO-NixOS


I particularly enjoy the Asterix-adjacent naming scheme.

I’m glad multiple European agencies are giving this a whack - hopefully we’ll end up with a synthesis of the best aspects.

[flagged]


Does it? The EU isn't a single country like the US - it's a partnership of 27 separate countries.

If one country in the UN adopted some new OS, we wouldn't expect every other country (incl the US) to adopt it too.


That's an accurate description of the status quo, but it's also a good example of massive inefficiencies in the EU.

FWIW, I'm not saying that I would expect the EU to do a better job than the Dutch government. Likely quite the opposite. But it's frustrating as passionate European... that these problems aren't tackled on a union level. You'd think that's their job? The EU will never compete if it stays a loose partnership of 27 countries, and if it can't even get France, Germany, and the Netherlands to work on this together instead of running 3 separate projects...


This is the equivalent of two competing startups getting an advantage over an enterprise, just at government scale. It's a good thing.

Two competing products emerge, they are battle tested, opinions form, reviews are conducted. At that point it's a much better time for a significantly larger organisation to create a more focused realistic plan based on learnings and outcomes.

Imagine the analysis paralysis that would set in if the EU tried it from scratch - years of requirements gathering, pilots, reports...


> it says a lot about the EU that every government tries to reinvent the same operating system

Yeah what could possibly go wrong


[flagged]


Not all of us know, could you give me a hint please? Why is NixOS toxic?

Ask your llm, scroll through some meme pages!

https://x.com/MemesOfNixOS/status/1942271466894217321


- Could you please justify your opinion? - Sure, here you have a link from Twitter And we all get a meme that is not funny.


Can an OS be toxic?

From my recalling, a few years at most. There was, and apparently still exists, distributed.net which was aimed at brute forcing DES (easy), RC5-56 bits and then RC5-64 bits by establishing a web of personal computers (via a client one had to install). Thus it was well known brute forcing was achievable in a reasonable time.

PGP (1991) was considered secure as it was considered not brute forceable. With 128 bits, it was considered military grade at the time and the US had an export restriction due to that. That might have been an incentive for GNU Privacy Guard. In France you had to give your private key to the government authority if an encryption system used anymore than 56 bits (as I recall, I don't remember the exact number).


Copyright IS an issue to Debian, and always has been for the last 30 years or so. They are, rightfully so, extremely picky when it comes to respecting copyright and licensing. It is the 0.01%.

> Who the f*ck cares who made it? A monkey could have made it for all I care. If it does what it claims to do, and I understand how, it's all good.

Copyright laws do care. As an example one can send a patch claiming its their own but because they had do it under their employer duty, the copyright might well be associated to their employer rather than them individually. Does the patch does what it claims to? Surely. Is that a copyright infringement? DEFINITELY SO.

The copyright rationale for the first proposition (Choice 1: Ban LLM contributions from Debian via Social Contract):

> 1. Copyright

> -------------

>

> LLM output has very unclear legal status: it may be possible to copyright on its own merits, or not; it may be affected by all of the licenses and copyrights in the training data, or not.

> Debian Policy and the DFSG require absolute clarity for licensing and copyright[1][2]. Software and other contributions written conventionally by humans with unclear copyright or license status are not allowed in Debian; LLM output should not have a special exception to this.

This rationale states if there is doubt about the copyright of the code, it not suitable for inclusion. Until I guess LLM output get a clarification regarding who is the author of its output.


Maybe but I think you are underestimating the achievements Fabrice has accomplished. Among others: - Improved an algorithm to compute Pi, ran it on a *personal laptop* and broke the world record. That achievement is not even listed on his personal homepage, and it a single line of facts with Zero bragging involved https://www.bellard.org/pi/pi2700e9/ - a PC emulator in vanilla javascript, boot the Linux Kernel in a browser and get a virtual terminal also implemented from scratch - QuickJS, embeddable, self contained (no libs) and fast JavaScript engine matching almost entirely ES2025 - NNCP, a Neural Networks driven lossless data compression system

And more https://www.bellard.org/

I have been referring to his page for decades as an example of one can have a huge respect without having a fancy web page and no bragging at all. He is a genius :-)


Not at all. I mean, regardless of him not having a fancy web page or an Instagram, he is anyway an Internet geek celebrity we all know and respect. My point is that I believe there are many similar but noname engineers whose achievements stayed and will stay behind corporate proprietary walls.


I still use « briques », typically « 10 briques » instead of « 100 k ». I think there is some poetry in sticking to the old obsolete term.


I do not understand why the scrappers do not do it in a smarter way: clone the repositories and fetches from there on a daily or so basis. I have witnessed one going through every single blame and log links across all branches and redoing it every few hours! It sounds like they did not even tried to optimize their scrappers.


> I do not understand why the scrappers do not do it in a smarter way

If you mean scrapers in terms of the bots, it is because they are basically scraping web content via HTTP(S) generally, without specific optimisations using other protocols at all. Depending on the use case intended for the model being trained, your content might not matter at all, but it is easier just to collect it and let it be useless than to optimise it away⁰. For models where your code in git repos is going to be significant for the end use, the web scraping generally proves to be sufficient so any push to write specific optimisations for bots for git repos would come from academic interest rather than an actual need.

If you mean scrapers in terms of the people using them, they are largely akin to “script kiddies” just running someone else's scraper to populate their model.

If by scrapers in terms of people writing them, then the fact that just web scraping is sufficient as mentioned above is likely the significant factor.

> why the scrappers do not do it in a smarter way

A lot of the behaviours seen are easier to reason if you stop considering scrapers (the people using scraper bots) to be intelligent, respectful, caring, people who might give a damn about the network as a whole, or who might care about doing things optimally. Things make more sense if you consider them to be in the same bucket as spammers, who are out for a quick lazy gain for themselves and don't care, or even have the foresight to realise, how much it might inconvenience¹ anyone else.

----

[0] the fact this load might be inconvenient to you is immaterial to the scraper

[1] The ones that do realise that they might cause an inconvenience usually take the view that it is only a small one, and how can the inconvenience little them are imposing really be that significant? They don't think the extra step of considering how many people like them are out there thinking the same. Or they think if other people are doing it, what is the harm in just one more? Or they just take the view “why should I care if getting what I want inconveniences anyone else?”.


Because that kind of optimization takes effort. And a lot of it.

Recognize that a website is a Git repo web interface. Invoke elaborate Git-specific logic. Get the repo link, git clone it, process cloned data, mark for re-indexing, and then keep re-indexing the site itself but only for things that aren't included in the repo itself - like issues and pull request messages.

The scrapers that are designed with effort usually aren't the ones webmasters end up complaining about. The ones that go for quantity over quality are the worst offenders. AI inference-time data intake with no caching whatsoever is the second worst offender.


Because they don't have any reason to give any shits. 90% of their collected data is probably completely useless, but they don't have any incentive to stop collecting useless data, since their compute and bandwidth is completely free (someone else pays for it).

They don't even use the Wikipedia dumps. They're extremely stupid.

Actually there's not even any evidence they have anything to do with AI. They could be one of the many organisations trying to shut down the free exchange of knowledge, without collecting anything.


The way most scrapers work (I've written plenty of them) is that you just basically get the page and all the links and just drill down.


So the easiest strategy to hamper them if you know you're serving a page to an AI bot is simply to take all the hyperlinks off the page...?

That doesn't even sound all that bad if you happen to catch a human. You could even tell them pretty explicitly with a banner that they were browsing the site in no-links mode for AI bots. Put one link to an FAQ page in the banner since that at least is easily cached


When I used to build these scrapers for people, I would usually pretend to be a browser. This normally meant changing the UA and making the headers look like a read browser. Obviously more advanced techniques of bot detection technique would fail.

Failing that I would use Chrome / Phantom JS or similar to browse the page in a real headless browser.


I guess my point is since it's a subtle interference that leaves the explicitly requested code/content fully intact you could just do it as a blanket measure for all non-authenticated users. The real benefit is that you don't need to hide that you're doing it or why...


You could add a feature kind of like "unlocked article sharing" where you can generate a token that lives in a cache so that if I'm logged in and I want to send you a link to a public page and I want the links to display for you, then I'd send you a sharing link that included a token good for, say, 50 page views with full hyperlink rendering. After that it just degrades to a page without hyperlinks again and you need someone with an account to generate you a new token (or to make an account yourself).

Surely someone would write a scraper to get around this, but it couldn't be a completely-plain https scraper, which in theory should help a lot.


I would build a little stoplight status dot into the page header. Red if you're fully untrusted. Yellow if you're semi-trusted by a token, and it shows you the status of the token, e.g. the number of requests remaining on it. Green if you're logged in or on a trusted subnet or something. The status widget would links to all the relevant docs about the trust system. No attempt would be made to hide the workings of the trust system.


And obviously, you need things fast, so you parallelize a bunch!


I was collecting UK bank account sort code numbers (to a buy a database at the time costs a huge amount of money). I had spent a bunch of time using asyncio to speed up scraping and wondered why it was going so slow, I had left Fiddler profiling in the background.


« mann » comes from Old English and stands for a human being.

« wïfmann », literally "female human", led to « wife ».

« were » means man and comes from Germanic and I don't think « weremann » has ever been a thing.


Blew my mind when I found out “world” is a direct derivative of the word “were”.


“world” is derived from “were” + “eald” (old), and meant “the age of humans”, which was distinguished from the age of the Gods, when the Æsir and Vanir dominated, and the age of the Jötnar.

I find it interesting how the term shifted from a (mythical) temporal concept to a spatial concept, to now often a social concept (e.g. the Fourth World).


Why add the complexity of having to maintain an Ansible installation, a logging stack, deal with their upgrades and whatever python issue one might encounter. I had the issue of Ansible builtin `shell` not doing the right thing (sh vs bash) or it being unnecessarily slow when uselessly looking up `cowsay`.

Adding layers and layers of tooling is often overkill and it is hard to bit the simplicity of 33 lines of shell when the use case is a single person doing the code, deployment and maintenance.


I’m with you on the usecase. Simple server deployment on a VM, bash script is fine, in fact I recommend it. It’s when you start dealing with 5+ VMs that I would start looking into using a tool like Ansible.


@unixispower , you might consider adding the site to TheOldNet webring ( https://webring.theoldnet.com/submit ). I have discovered it from a yesterday post about ucanet (a DNS for retro site). Your site would be an excellent ring member!


Thanks for the heads up. I'll have to make up some banners later and submit. I have a few sites that would go nicely there.


> inability to moderate which is really important for adult content

Given Elon Musk twitted about moderation being censorship (twist: it is not), what could go wrong!?!


His position was about moderating legal content. He thinks moderating legal content is censorship. He is in favor of taking down illegal content. If people think currently legal content should be taken down they should appeal to change the laws and it shouldn't be part of at least their platform to judge the legal content. That's his position, not my position.


That's what he says his position is. He bans people who post the movements of his private jet. He also has had Twitter file lawsuits against people merely for saying mean things about Twitter.

Of course, it's also possible that he's a complete idiot and has no idea what the First Amendment actually permits in speech, like he stopped paying attention after Schenck v US.


Just because he's a hypocrite, doesn't mean we should censor more.


Yeah, that is hypocritical. He should be congruent in his views and in any ambiguity in content disfavoring him he should lean towards giving benefit of the doubt to prove lack of bias. Although in this case there didn't seem to be any ambiguity, he should have not banned the account.


Filing lawsuits is entirely consistent with a position that the courts, not private companies, should regulate speech online.

Then there's Jack Sweeney, the guy tracking Elon and his private jet - who, by the way, was accused of facilitating stalking by Taylor Swift. For him, X made a policy against any account "doxxing real-time location info of anyone".

Does anyone here actually want to argue that tracking real-time location info should be allowed?


I'll bite.

Real-time tracking of a person or their ground transport: not ok.

Real-time tracking of a person's plane: ok.

And the reason for such a distinction is that planes can only land on specific ground slots. That also means that real-time tracking of a person's helicopter falls under "not ok". And by extension, the same will hold for flying taxis, once they take off[tm].

We don't (yet?) live in a world where shoulder mounted surface-to-air missiles, outside of war zones, are a realistic threat.


Once a stalker knows a specific place and time to find their victim, they can simply follow them until they have an opportunity to do worse.

"Missiles" are just a straw man.

Hopefully Taylor Swift will follow through, Sweeney will be sued or even prosecuted for stalking, and we'll find out if this really is legal.


He just banned posting the identity of pseudonymous accounts which is definitely not illegal. He also banned posting public information about the movements of his private jet which is also definitely not illegal.

This stuff would easier to take seriously he was consistent about it. At this point it’s kind of insulting


This is one where “won’t someone thing if the children” (underage and non-consensual) is relevant. It tends to be a lot more… universally agreed upon limits to whatever you think free speech is.


> Elon Musk twitted about moderation being censorship

And instead, X has implemented hellbanning, where nobody outside X knows who is being censored and why. People just slowly figure out that, actually, no one sees their posts. At least with outright <scare-quote> "moderation" you would know that you had been cancelled.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: