Spyke

Syndicated from the fediverse. Read and engage on the original instance.

View original on thebrainbin.org
nostupidquestions·No Stupid QuestionsbyInfrapink

How does reader mode work?

I tried googling and I got a lot of hits for what reader mode does and how to enable it, but nobody can explain what the browser is actually doing.

Do sites have pared-down versions of their pages specifically for reader mode? Or does the browser just scan for particular HTML tags?

View original on thebrainbin.org
43

7 replies

A web page is a combination of a structured document and a separate stylesheet. The same structured document can be styled in radically different ways just by changing the stylesheet. This is how a website can offer different skins, light/dark mode, etc.

Consider this basic HTML document:

<html>
   <head><title>Test Page</title></head>
   <body>
      <p>This is a paragraph.</p>
      <ul>
         <li>This is an unordered bullet-point list item</li>
         <li>This is also a list item</li>
      </ul>
   </body>
</html>

And here's a stylesheet to make the bullet points square:

ul {
  list-style-type: square;
}

Change the stylesheet and you change how the document is rendered. "Reader mode" is just a minimalist stylesheet built into the browser.

27

I think OP is asking how the browser figures out what is the important content to display in reader mode and what's the cruft that can be dropped.

9

If it helps there's a standalone version of the logic.

At a rough scan, it looks like it tries to find a best guess "main content" node by stripping unlikely nodes and then scoring each node. Element type and content contribute to the score

/**
       * Loop through all paragraphs, and assign a score to them based on how content-y they look.
       * Then add their score to their parent node.
       *
       * A score is determined by things like number of commas, class names, etc. Maybe eventually link density.
       **/
15

I don't have any insider knowledge, but I'd guess they're basically just rendering the html without JS or the original CSS, probably with a few tweaks. They may also look for certain tags (like "content") to prioritize the main section. Web developers have a vested interest in notating the content for SEO reasons.

6

It's both. The browsers can do it automatically by detecting the content from the page, but the site can also use specific HTML as a sort of instructions for the browser m.

3

I believe this is due to ADA Title III requirements for publicly accessible businesses:

The ADA requires that businesses open to the public provide full and equal enjoyment of their goods, services, facilities, privileges, advantages, or accommodations to people with disabilities. Businesses open to the public must take steps to provide appropriate communication aids and services (often called “auxiliary aids and services”) where necessary to make sure they effectively communicate with individuals with disabilities.

https://www.ada.gov/resources/web-guidance/

Screen readers and braille displays can't do anything with all the fancy JavaScript and other nonsense that frames and formats the content for delivery, including the user-identifying-and-blocking stuff. The website is legally required to make the content accessible in a way that assistive technologies can use, and usually the simplest way to do that is to just deliver plain text. I think "reader mode" is leveraging this functionality to get the content in a simple form without the cruft, basically just telling the website that the user needs the ADA compliant version.

2

You reached the end