Plain text is still one of the best technologies we have

(deadparrotbbs.com)

158 points | by speckx 2 hours ago

26 comments

  • devy 1 hour ago
    Graydon Hoare, the creator of the Rust programming language, wrote his seminal piece on text in 2014. In it, he said "text is the most powerful, useful, effective communication technology ever, period."[1] Text is durable.

    [1] https://archive.ph/FhG5L (the original either got deleted or login-walled, here is the archived version)

    [2] https://hn.algolia.com/?q=always+bet+on+text

    [3] https://news.ycombinator.com/item?id=26164001

    [4] https://news.ycombinator.com/item?id=8451271

    [5] https://news.ycombinator.com/item?id=10284202

    [6] https://news.ycombinator.com/item?id=12815829

  • boomlinde 43 minutes ago
    The article briefly addresses the problem, but it's pretty fun how different "plain text" looks throughout history and in different domains.

    For one, there are a few different ways to terminate lines. All major operating systems now tend to use just \n, but I have older files that use \r\n (Microsoft), \r (Macintosh) or \n\r (RiscOS).

    There are also different opinions on how text files end in different operating systems. In POSIX, all lines are terminated by \n, even the last one. Microsoft software still tends to insist that the last line of a file is special case that doesn't need to be terminated even now that they have otherwise adopted POSIX style line endings. In Microsoft's view, it seems that the line ending sequence separates lines rather than terminate them. Files created according to this view don't play well with tools like cat(1) if your intent is to concatenate the lines of two files, but it seems other Unix clone tools have adapted to the possibility that the last line isn't terminated properly.

    Finally there's the encoding problem. I don't know of a good tool that determines the original encoding based on heuristics and re-encodes to UTF-8 but if someone does I'd love to know. If I know that the input language is English for example it shouldn't be too hard to determine what encoding the funny byte used in contractions or the funny bytes used in quotes belong to. Still, in English most of the files that use 8-bit encodings remain quite readable if you just box out the invalid bytes.

    • adiabatichottub 18 minutes ago
      Text is wonderful, but note that the one keyword not found in this text is 'parse'. How much extra work is generated writing parers for text that would have been so much easier to deal with if it had just been generated in a structured binary format? You still get the pleasure of designing a new wheel with every format.
    • Liquid_Fire 26 minutes ago
      Sorry, can you clarify about Microsoft having adopted \n? This is the first time I hear of it, and I can't find anything online about it.
    • gwbas1c 36 minutes ago
      This morning I had co-pilot generate a LINQPad script that fixes a lot of that.

      Granted, I knew that all files are UTF-8, so I didn't need to have it "guess" what the encoding was without the BOM.

  • drhagen 1 hour ago
    This is the public dividend of a standard finally winning. The article gives credit to Unicode, but it is the fact that ASCII unambiguously won that gives plain text its portability and longevity. It looks like Unicode is on its way to winning in the same way, but it is not there yet. Most text files I write are still pure ASCII because that's the only way to avoid unexpected glitches [0].

    [0] Windows newlines not withstanding.

    • bruce511 1 hour ago
      Ahh, plain text, wherein "plain" does some heavy lifting. If plain text had something as simple as a 4 byte signature it could have been soooo much better.

      As it is programs have to guess the following;

      A) encoding. If ANSI which code page? If unicode which encoding? In the case of utf-16 big or little endian?

      B) line endings? CR? LF? CRLF? I guess no-one uses LFCR... right?

      C) number formats? 100.000 or 100,000? Date formats? is that mm-dd-yyyy or dd-mm-yyyy?

      D) what human-language is it in?

      E) CSV? Don't get me started...

      Yes. Text is a lot easier to load, parse, make guesses about than say XLS. Yes all of the above things can be "guessed" to a greater or lesser extent.

      Yes BOM exists to at least try solving the encoding question. Of course most files don't use them. Lots of tools don't support them. And they don't solve any of the other issues.

      But sure, ASCII files in English with US date formats, and Windows line endings....no problems at all...

      • hnlmorg 44 minutes ago
        ASCII is just a raw binary format. It isn't just for text.

        The line endings problem really isn't a problem. Pretty much every text editor out there can handle different line endings.

        I don’t think number nor date formats are relevant here. For example you could have that same problem entering text into MS Word. That’s really more of an issue if you want to use text as a database rather than a document format, which isn’t really something that even plain text advocates would generally recommend.

        As for human language detection, that’s a much easier problem to solve than decoding a proprietary binary blob.

        CSV definitely has its warts. But it’s not like that’s the only plain text option for serialising data. (JSON, jsonlines, YAML, XML, etc). Or you could use the actual ASCII codes reserved for records, if you really wanted something that didn’t require quoting and escaping in plain text. It’s actually a pity nobody does this.

      • conartist6 47 minutes ago
        I've made a language called CSTML to do some of this: https://docs.bablr.org/guides/cstml

        Basically it's a text-based language for embedding semantic metadata into some kind of underlying text stream. Is this interesting to you?

      • jimbokun 1 hour ago
        [flagged]
    • layer8 50 minutes ago
      > Windows newlines not withstanding

      “Windows” newlines are also the standard in many communication protocols like HTTP and SMTP. That’s not because of Windows or DOS, it’s because it was the standard for teletypes which needed bot CR and LF. It’s arguably systems like Unix that deviated from that standard.

      I agree that beyond ASCII there is a slope from well-supported to less-supported and quirky to problematic areas of Unicode. For example, HN filters many Unicode text elements like combining characters (Zalgo text) and emojis, and there is no specification to point at what it supports.

      Even within ASCII, most control characters don’t have a portable meaning. So it’s really just the printable subset of ASCII, and strictly speaking not even that, given that there are regional variants of ASCII, such as the Japanese one where backslash becomes the Yen sign.

      • hnlmorg 39 minutes ago
        UNIX (and Linux) is even more annoying because pseudo TTYs will require CRLF when in raw mode but requires only LF when in line mode.
    • tyromaniac 1 hour ago
      And file endings..
    • Analemma_ 1 hour ago
      If you write files in ASCII you’re already writing in UTF-8.
      • sillysaurusx 1 hour ago
        That’s true, though if you’re writing UTF-8 you have to handle the corner cases when reading UTF-8, of which there are many.

        Fortunately they’re easy to test for and most languages have standard libraries that make this painless.

      • bell-cot 1 hour ago
        ASCII has, in principal, infinitely many superset.

        And in practice, still a rather large number of them, going back to the 1970's - https://en.wikipedia.org/wiki/Extended_ASCII

  • andsoitis 21 minutes ago
    The more easily you can access meaning from symbols without intermediate tools, the more durable and resilient.

    Symbols on paper is best.

    Plain text in a file that can be opened by any computer comes close, but needs a tool.

    Fancy file formats that need not only hardware but also special software are the worst.

    However, there's an important tradeoff, which is the fancier formats can present information in ways mere symbols might struggle with, and can also use interactivity to improve understanding.

  • kerblang 1 hour ago
    What, no mention of our old friend, ASCII-armored Base64? For shame! Everything can be text with Base64, including things that have absolutely no business being text! Best of all, in light of popular widespread abuse of every available resource, Base64 is comparatively efficient! Bring on the petabytes! Yay and I'm not being completely sarcastic
    • adiabatichottub 50 minutes ago
      And then you add RFC2045 headers so you know what the encoding is...
      • kerblang 47 minutes ago
        Text headers, mind you! Ah, text text text

        Don't ask me what my point is

  • spcebar 1 hour ago
    Fun timing on this plaintext conversation. I built a little browser based plaintext playwriting app this weekend for a little weekend project. Found myself really hating 1. how clunk screenwriting software can be and 2. How unsharable the files are.

    With plaintext you can hand the file off to anyone on any device--the caveat being absolutely no one wants to be handed a plaintext script. The software I used ten years ago to write plays is long since deprecated and those files are basically unopenable. Plaintext however remains.

    • someonebaggy 39 minutes ago
      When you're looking to make custom markup, XML is also a pretty good base. Ignore all the silly stuff like DTDs and namespaces, and just use the tags as tags and the text as text.
      • spcebar 13 minutes ago
        XML is a good candidate, but actually overkill for the purposes I was looking to achieve. Basically, all I needed was character names are all caps, scene headings have hashes, and stage directions are free text. Put a handful of other niceties like title pages, casts, and page breaks, but never anything so complex it needed additional markup.

        # INT. JOE'S BEDROOM, DAY.

        JOE

        This is my dialogue, yeehaw.

        JOE exits, and we know JOE is exiting because this is a stage direction, and we know this is a stage direction because it is free text and doesn't satisfy any other formatting requirement.

    • dunham 1 hour ago
      John August has also talked about using plain text for movie/television screenplays, specifying a format based on markdown and hollywood standard: https://fountain.io/

      Sharing aside, I imagine a plain text format would also be helpful for version control.

      • spcebar 34 minutes ago
        Haha, yes, as I was finishing writing the parser, I stumbled upon Fountain and realized my genius idea for a markdown based playwriting format was not so original. Ended up adding some Fountain syntax to my parser and ultimately I built a very functional but not feature complete implementation of Fountain by accident.

        I built sharing and (sort of) versioning into the web app. Everything is saved in localStorage and on a MySQL server. You can Save or Save As, and every time you Save As it inserts a new database row and gives you a new, sharable URL.

        I wanted an experience where you didn't have to login to use it, but ultimately, having a way to keep track of all those URLs would actually be a pretty decent versioning system--add diff checking between script versions and you end up with something pretty handy.

  • mythrilkey30 10 minutes ago
    I dropped the first 3 paragraphs into pangram and it is 100% AI generated. Not reading and moving on.
  • shmolyneaux 1 hour ago
    I love plain text, but the calculus of being able to access the content in 50 years is much more interesting in the age of agents. They can infer the meaning of structured data without schemas, decompressed archives, find embedded files, etc. It's not perfect, but the durability of binary formats is better now than it has ever been.

    It's not baked-in to the weights of any model, but agents can write their own tools to work with arbitrary binary formats and get many of the benefits of off-the-shelf unix utilities.

    I love that the Godot game engine has a textual scene description. That's a stark contrast to Unreal Engine's binary format for blueprints (visual scripting).

  • frollogaston 40 minutes ago
    I greatly prefer plaintext in most cases where Markdown ends up being used. WYSIWYG is more useful than formatting in these cases. Most of the formatting makes it harder to read anyway. Some say you can ignore MD and treat it like plaintext... not so when you use a newline.

    Oh and now you could have LLMs write the Markdown, but chances are no human reads that, or even if you do, the formatting is going to be insane. I have to keep telling Claude to give me a .txt instead of .md when asking for a context dump. Maybe .md is popular for agent readmes because they optimize around its headings.

  • CrimsonCape 42 minutes ago
    Hijacking the plain text discussion to ask what the best tools are to convert plain text to lexed/parsed output?

    I'm assuming any tool in this regard would expect the user to write an EBNF grammar.

    I found ANTLR to be nice but it's way too convoluted to use as a tool with Java dependencies and seems to be stagnating. And tree sitter just is too convoluted and requires the added C/C++ overhead to understand how to use it.

  • njarboe 52 minutes ago
    I work on scientific data repositories [1] and pushed for plain text files as our archiving file format two decades ago as the best long term format. We implemented it on our data repositories. Very human readable also. The datasets are small and table based, so this works well. We latter started archiving 2-D image data and that creates very large files. We are still looking for an elegant solutions for those datasets.

    [1] https://earthref.org/FIESTA/

  • pratikdeoghare 45 minutes ago
    I like plain text. I tolerate markdown. Markdown was designed as a shorthand for html. It gets stretched to do other stuff.

    I created Brashtag [1]. It is simpler than markdown.

    [1] https://github.com/PratikDeoghare/brashtag

    • CrimsonCape 31 minutes ago
      Can you think a little more about your decision to strongly type code blocks. Some domains have no need for code escaping but do need blob text escaping (which you already have defined elsewhere.) Technically you could treat code via 'this is a bag named code with a blob, therefore it is code."
      • pratikdeoghare 2 minutes ago
        I wanted to keep the famaliar things. We write code blocks with backticks in markdown and slack and everywhere so I used those. In fact I used $ signs earliar then switched to backticks.

        > 'this is a bag named code with a blob, therefore it is code."

        This would make #code{} a special bag. Right now no bag is special. Also it would be hard to find the closing } if somebody wrote unbalanced paren code in the bag like #code{ func main() { }.

    • someonebaggy 40 minutes ago
      I just use HTML if I want HTML. No guessing what formatting characters mean - <b> means bold and <h1> means heading 1, period.
      • pratikdeoghare 8 minutes ago
        What if you wanted say footnotes? html does not have it.

        Imagine you could just say #footnote{This will become footnote} and have a little program that would produce html that will show it at the bottom of the page.

    • gandreani 42 minutes ago
      Why did you create it? What were you looking to solve address?
      • pratikdeoghare 14 minutes ago
        I had multiple frustrating experiences. I kept thinking I could just write things in some format in text and have some programs process it to do different things. So I created it.

        See how easy it is to write programs to process brashtag [1].

        I can chat with llm, interact with jupyter kernel and do literate programming from any text editor all in the same document [3].

        It took me like 10 minutes to build this mermaid like thing [2].

        List of frustrations:

        - One day Anki crashed. --- You can just write #card{....} in a text file. Some program would read it and show flashcards in browser.

        - Another day JabRef crashed. --- You can just write #bib{....} put link to a paper and its citation info there and have a program download the paper and copy the citation to bib file.

        - Markdown parser processed mathjax wrong. Why keep trying different markdown parsers? Just write #h1{}, #b{} #i{} and just convert them to html tags.

        - Literate programming tool I was using failed at some corner case.

        - My notes were scattered across files and tools. I couldn't find anything. I tried to have a text file where sections were separated by four dashes. But that didn't allow nesting. Also, the parser I wrote for that hit corner cases. Why not just put #note{} #JIRA-334{} etc. in text file.

        [1] https://github.com/PratikDeoghare/brashtag#some-programs [2] https://github.com/PratikDeoghare/brashtag/tree/master/cmd/m... [3] https://www.youtube.com/watch?v=IMXgIE0Vljg

  • anvuong 58 minutes ago
    ASCII was amazingly efficient for conveying English text. Then we needed to encode multi languages and emojis, the resulting Unicode is just a mess.
    • Perseids 32 minutes ago
      Human languages are a mess. Some mix left-to-right and right-to-left. Some have variable amount of diacritics. Some have non-well-defined character sets where just trying to map it to a manageable complexity still loses cultural heritage. When you are fondly remembering the good old times of ASCII, you just wish back to be able to ignore anyone not speaking English. (Which is in your right to do, if you so desire.)
    • weinzierl 49 minutes ago
      Unicode was mostly OK when it wanted to encode multi languages. It started to get a mess when thought it needed to do more.
      • frollogaston 37 minutes ago
        Yeah the multi-unicode-char grapheme clusters are basically only for emojis once they started adding stuff like every ethnicity/gender combo of 4-person family. And that one historical Korean script.
      • OkayPhysicist 31 minutes ago
        What do you think doesn't belong in Unicode? It turns out that expressing all the different ways humans have used text in history has some necessary complexity, but assigning each character and modifier a number seems like a perfectly reasonable approach.
        • weinzierl 11 minutes ago
          "have used text in history"

          That's how Unicode started but nowadays Unicode expresses things that are only used because Unicode itself introduced them.

          And this while many important things in the "have used text in history" category have been unfinished or not tackled at all.

  • Grimeton 1 hour ago
    ASCII (7-Bit) is the only widely understood charset there is. Everything beyond this point depends on the loaded charset.

    The fact that unicode maps the lower 7 bits to its own character set is a nice touch but none of the unicode sets are plain text.

    Unicode are multibyte characters with variable byte length and endianess at play. If you read it wrong or guess the length wrong your results might be anything but useful.

    • Perseids 47 minutes ago
      You seemed to be deeply confused about encodings and character sets.

      > The fact that unicode maps the lower 7 bits to its own character set is a nice touch but none of the unicode sets are plain text.

      That is only true for the mapping of Unicode character (code point to be exact) to UTF-8, which is an encoding of Unicode characters.

      > Unicode are multibyte characters with variable byte length and endianess at play.

      That is only true of the UTF-16 encoding. UTF-8 does not have endianess, UTF-32 does not have variable length per code point.

      None of that is true for Unicode, because it is abstracted away from any byte representation. Furthermore, getting back to the start:

      > ASCII (7-Bit) is the only widely understood charset there is. Everything beyond this point depends on the loaded charset.

      ASCII also depends on how you try to decode your text. If you interpret two bytes as one character, ASCII will never be correct. There is nothing magical about one byte mapping to one character. I'd even argue that the only reason ASCII support is so universal internationally is because of UTF-8. Otherwise many countries would default to encodings where ASCII is not a subset (as they did before UTF-8 became common). So IMHO, UTF-8 and Unicode are the only widely understood encoding and character set.

  • vatsachak 1 hour ago
    Probably because plain text can encode any form of distilled data. Technically our DNA can be plain text lol
  • ollien 1 hour ago
    Tangential to the actual point of the post, but the talk "Plain text? Really?" by Dylan Beattie[1] is one of my favorite talks. It does a great job capturing the problems with something that "seems" so simple. I think the author of the blog is well aware of these, though :)

    [1] https://www.youtube.com/watch?v=_mZBa3sqTrI

    • s1mon 1 hour ago
      Many things I knew, and a bunch I didn't. Thanks for that.

      I wish he'd spent a little more time on the fundamental issue of CR vs CR/LF vs LF. That's been a "plain" text nightmare since before Unicode and many of the other complexities existed.

  • leftnode 38 minutes ago
    Because no single person or entity can own it.
  • dimiprasakis 1 hour ago
    Prediction: RSS is coming back stronger than ever
    • slowin 1 hour ago
      RSS was removed from most sites (unfortunately) due to business, not technical reasons. Just like Web 2.0 era "mashup" friendly APIs, business started locking down their data despite it being an extremely user-hostile move.
    • DonHopkins 1 hour ago
      .plan is really simpler syndication.
  • olexsmir 1 hour ago
    I feel obligated to mention ledger[0] and hledger[1], those are plain text accounting software, well, as you can guess from their names, they allow you to do personal accounting in plain text.

    0: https://ledger-cli.org

    1: https://hledger.org

  • mythrilkey30 9 minutes ago
    I pasted the first 3 paragraphs of this into pangram and it's 100% AI generated. Not reading, moving on. We really shouldn't be promoting slop on hacker news
  • m463 1 hour ago
    I think text is getting a big boost from the age of ai.
  • ummonk 48 minutes ago
    > The Unicode Standard defines plain text essentially as a sequence of character codes, without the additional formatting information associated with rich text. Fonts, colors, layout, and similar presentation details belong somewhere else.

    Nope. Skintones are part of Unicode. It also has 33 control characters from ASCII including one that rings a bell... There are also numerous characters added by Unicode that are literally called "layout controls".

    If you want a format that provides purely semantic information, then "plaintext" doesn't fit the bill.

  • jaekwon 1 hour ago
    that's why gno.land contracts render to markdown as the standard. you can browse the world through your terminal.

    imagine the world wide web but markdown (and decentralized). gno.land is that.

  • charcircuit 38 minutes ago
    >There are not many computer file formats I would trust to still be readable fifty years from now, but plain text is one of them.

    Just the other day I had opus read a 10 year old proprietary file format. The idea that we will lose the ability to use file formats is not consistent with reality.

  • Koshkin 1 hour ago
    "A picture is worth a thousand words"
    • smalltorch 1 hour ago
      There was a challenge awhile back to describe a picture in 1000 words can't remember where I saw it. The winner just made a encoder/decoder to turn jpeg into words. The image was somewhat compressed but contained more visual data then you could accurately represent in the same way a raw description could. I thought it was so cool.
    • mohamedkoubaa 1 hour ago
      Sometimes I'd rather have ten unambiguous words.
  • FLeXMurphy 1 hour ago
    @dang @tomhow

    My understanding is that HN has started incorporating AI tooling in the back-end to scan for LLM-generated submissions and source content, in an effort to discourage it and encourage human-made content (with human discussions, one would hope). Why do we keep having these kinds of articles every day? Entire "apps" are generated - see the lighthouse one - and submitted, and are plain-as-day LLM-generated nonsense.

    Anyway, flagged.

    • bityard 22 minutes ago
      I was about to argue with you, but I now believe this article (and indeed the whole blog) is AI-generated. Here's my evidence:

      * All the posts are walls of text. In many of the posts, all of the paragraphs are just about exactly the same length.

      * None of the posts contain personal stories, experiences, or anecdotes.

      * All of the posts are hyper-focused around advocacy of "old tech" and decentralization. Which is great, but real blogs with this many posts have at least _some_ variation in subject matter and quality. This one's surprisingly uniform.

      * The blog is relatively new, posts began about two months ago and there are 17 posts, that's a little over 2 posts a week. Sure, there are bloggers who post that often, sometimes daily, but they are usually much shorter posts, tend to be shorter, or do it for their job.

      * The picture of the author on the About page is very obviously AI generated.

      * The author has a GitHub account that had virtually no activity until March of this year, and basically all of the commits were co-authored by Claude.

      To be clear, I'm sure a real human is behind the site/posts (as opposed to a fully autonomous agent), but I'd bet dollars to donuts that each article started out its life as a very short prompt.

    • sillysaurusx 1 hour ago
      Where did you get that understanding from? I might have missed something. They penalize LLM comments. Do they also penalize submissions? Any evidence of that?
      • FLeXMurphy 1 hour ago
        dang mentioned it in passing some weeks back (going off of memory). Also please publish new NoH/Kongor build.
    • sthatipamala 1 hour ago
      I skimmed this whole article and didn't see any classic LLM tells about it. Why do you think it is LLM generated?
      • tux3 1 hour ago
        The recently released models have improved their writing styles. I don't know about this article, but if that's any indication of their writing habits, the author's github pages repo is full of posts generated by Claude.
        • pixl97 58 minutes ago
          Here's the thing, HN'ers have detected 200 of the last 10 AI written articles posted. We really seem to suck at it.

          Worse the tools for detecting it have an insane false detection rate. ESL = FP. Good writer = FP. Garbage human slop writer = A'ok.

          • tux3 24 minutes ago
            Detection tools in general have a bad selectivity and sensitivity, but there's consensus that Pangram is generally accurate and reliable. I don't use these tools much, but that's the word among people I trust, and that's what I see when I try it.

            And, well, neither of us really knows the ground truth. Maybe the guy who has a repo filled with articles co-authored by Claude is actually not posting slop this time and HN users are wrong and Pangram flagging it is just a false positive. But I can't blame anyone for being tired of giving this the benefit of the doubt.

            I'm ESL. The mistakes I make are part of what makes me human. If HN'ers are having an allergic reaction at the first hint of slop, it's because we've been drowning in it for months.

      • gwbas1c 39 minutes ago
        IMO, the article is well-written.

        The author's picture clearly has an LLM-generated background, (and is a little tacky, IMO.) That being said, I wouldn't dismiss an article merely because someone used an LLM while setting up their blog.

      • loup-vaillant 1 hour ago
        The whole structure had a faint smell to it. It's hard to put into word, there was no obvious "it's not the quiet part. It's the load bearing bedrock", but the way the thing flowed made me strongly suspect LLM involvement.
        • comradesmith 3 minutes ago
          First we had AI psychosis, now we have AI schizophrenia
    • john_strinlai 1 hour ago
      they don't get @ mentions, you're better off emailing hn@ycombinator.com
    • grebc 1 hour ago
      I think you’ve posted on an unrelated topic.
      • FLeXMurphy 1 hour ago
        No, this is the right article.
        • grebc 1 hour ago
          Doesn’t read like AI.

          Even has a simple an/a mistake that I doubt a chatbot is going to make.

    • thuruv 1 hour ago
      perused the posts and incidentally all are touching high-octane topics, posted only recently but surely warrants discussion/disputes, clearly an evidence of karma farming!
    • beej71 1 hour ago
      The is real is real.
      • FLeXMurphy 1 hour ago
        It is littered with little tells like that.