twenty eight of them · 14 min read · 31 aug 2026
list of things about the web you keep nodding along to
You can ship a whole site without knowing any of this. That is the problem. The thing works, so nobody explains why. Then something breaks and every word in the error message is a word you have heard but never checked.
These are in the order they stack. Each one only needs the ones above it.
Git and GitHub are not here. They sit one layer to the side, in list of tools that run a website, along with everything else you log into.
1. the internet
Physical machines connected by physical cables. That is the whole thing.
Some machines ask for things. Your laptop, your phone. These are clients. Some machines sit in buildings, switched on permanently, waiting to hand things over. These are servers.
There is no cloud. There is someone else's computer in a room in Virginia, and a cable.
2. a website
A folder of files sitting on one of those permanently-on machines.
That is genuinely all it is. The same kind of folder you have on your desktop. Files in it, some folders inside it. The only thing that makes it a website instead of a folder is that a machine somewhere will hand those files to anyone who asks for them.
3. DNS
Every server has a numeric address, like 76.76.21.21. Nobody can remember those.
DNS is a global phonebook that translates a name into the number. Someone types anusha.fyi, DNS returns the number, the browser goes to that number.
Buying a domain means renting a name in that phonebook and pointing it at a server. You are not buying a website. You are buying an entry in the directory.
4. what happens when you type an address
Five steps, in about 300 milliseconds:
- Browser asks DNS: what number is
anusha.fyi? - DNS answers with the number.
- Browser sends a request to that machine: give me the homepage.
- The machine sends back files.
- The browser reads the files and draws the page.
Every failure you will ever see is one of these five steps not happening. Knowing which one narrows the problem enormously.
5. HTML
The content and structure. The skeleton.
<h1>AI SDR Stack</h1>
<p>Compare 40 tools with real pricing.</p>
<button>See the list</button>
That renders a heading, a paragraph and a button. Ugly, but functional.
HTML alone is a complete website. Everything after this is optional.
6. CSS
The appearance. Paint and layout.
h1 { color: navy; font-size: 40px; }
button { background: black; border-radius: 8px; }
Same content, now it looks intentional. Remove the CSS and the site still works, it just goes plain.
7. JavaScript
The behaviour. Anything that reacts.
button.onclick = () => alert("Loading tools...");
Filters, dropdowns, search boxes, calculators, dark mode toggles.
The rule that makes all three stick: HTML is the noun, CSS is the adjective, JavaScript is the verb.
8. a page, a site, and a template
One page = one HTML file. about.html becomes yoursite.com/about.
A site with 78 tool listings could be 78 separate HTML files. Nobody writes those by hand. Instead you write one template, a single page design with blanks in it, plus a list of 78 tools in a spreadsheet or database. A build tool stamps the template 78 times, filling the blanks each time.
That is what "generating pages" means, and it is why a directory site can grow to thousands of pages without thousands of hours.
9. the folder, in full
A real project folder looks roughly like this:
my-site/
├── index.html the homepage
├── about.html
├── style.css
├── script.js
├── images/
│ └── logo.png
├── robots.txt instructions for crawlers
├── sitemap.xml list of all your URLs
└── package.json list of tools the project needs
Everything from images/ down is optional. The first four are the actual website.
10. why you cannot just email someone the folder
You could. They would have the files but no public address, and their copy would never update when you changed yours.
So you need two services:
- Somewhere to store the folder with a full history of every change.
- Somewhere to serve it to the public, permanently switched on.
Those are GitHub and Vercel respectively. They are not competitors. GitHub is the warehouse, Vercel is the shopfront, and Vercel reads from GitHub. Both are in the tools list.
11. publishing
Publishing is not uploading. That is the part that confuses everyone who learnt this in 2010.
The modern loop is:
edit files → save a snapshot with a note → send it to GitHub
→ Vercel notices → Vercel builds → live site updates
You never touch the server. You push to GitHub and the deployment happens because something is watching.
The build step is worth understanding. If the site is plain HTML there is nothing to build, the files are served as they are. If it is Next.js or React, the host runs a build command that converts your components into finished HTML, CSS and JS.
If the build fails, a typo or a missing file, the host stops and keeps the old version live. Broken code does not reach visitors. That is a real safety net and it is why this setup is worth the learning curve.
12. changing it
Same loop. Edit, snapshot, push, rebuild.
Two things make this safe rather than terrifying:
Preview deployments. Push to a branch that is not the main one and the host builds it anyway, at a private URL. Your redesign is viewable on a real phone before anything public changes. This is the single most useful feature in the whole stack and most people never use it.
Rollback. Every deployment is kept. If today's version is broken, open the dashboard, find yesterday's, click promote. Live in about ten seconds, no code involved.
You can always get back to a working site. Internalising that changes how much you are willing to break.
13. caching, and why your change is not showing
Your page is not served from one machine. It is copied to machines around the world so it loads fast wherever the visitor is. That network of copies is a CDN, and each copy is a cache.
The failure this creates: you push a change, the build succeeds, and the public still sees the old page, because the cached copies were never told to refresh.
You can check this. Every response carries headers, and two of them tell you the story:
age: 274030
x-vercel-cache: HIT
age is how many seconds old the served copy is. 274,030 seconds is 3.2 days. HIT means it came from cache rather than being generated fresh. A stale age with a HIT means visitors have been reading an old version of your site for three days.
Before you conclude a deploy failed, check whether it deployed and the cache simply never cleared. Those are different problems with different fixes.
14. crawlers
A crawler is a program that requests pages the same way your browser does, reads the HTML, finds every link on the page, and adds those links to a list of things to fetch next.
Google's is called Googlebot. It has been doing this continuously since 1998. Bing has one. So do ChatGPT (GPTBot), Claude (ClaudeBot) and Perplexity (PerplexityBot).
This is why links matter so much. A page with no links pointing at it and no sitemap entry is, from a crawler's perspective, invisible. It does not exist.
15. crawl, index, rank
Three separate stages that fail independently. Almost all confused SEO diagnosis comes from treating them as one thing.
Crawling. the bot fetches your page.
Indexing. having fetched it, the engine decides whether to keep it. It analyses the text, works out what the page is about, and files it. Plenty of crawled pages are never indexed: duplicates, near-empty pages, pages that add nothing to what is already stored.
Ranking. only now does a query enter the picture. Someone searches, the engine pulls every indexed page that could answer it, scores them, orders them. Under a second, and the ordering is different for every query.
Crawled and indexed are different states. Search Console shows you both and the distinction is where the diagnosis actually happens. A page that ranks nowhere might not be a ranking problem at all. It might never have been indexed.
16. robots.txt
A plain text file at your domain root that speaks to the crawl stage. A note at the front door listing rooms the bot should not enter.
User-agent: *
Disallow: /admin/
Sitemap: https://anusha.fyi/sitemap.xml
Two things everyone gets wrong.
It does not deindex. Blocking a page here stops the bot reading it, which can leave a stale entry sitting in results forever, because the bot can no longer see the page to learn it changed. To actually remove a page you use a noindex tag on the page itself, and you must leave it crawlable so the tag can be read.
It is public. Anyone can open yoursite.com/robots.txt. Never list anything you would rather people did not find.
You can also name specific AI crawlers here to block them:
User-agent: GPTBot
Disallow: /
17. sitemap.xml
A machine-readable list of every URL you want indexed. Speaks to discovery rather than permission.
<url><loc>https://anusha.fyi/list-of/north-goa</loc></url>
<url><loc>https://anusha.fyi/list-of/claude-code-thinking-words</loc></url>
It matters more for directories and archives than for blogs, because deep pages are often several clicks from the homepage and a crawler may simply never wander that far. You submit it once in Search Console and it gets re-read automatically.
Most frameworks generate it for you.
18. llms.txt
A proposed curated markdown index pointing LLMs at your best pages.
Worth being honest about this one. It is a community proposal, not a standard. An SE Ranking study of 300,000 domains found 10.13% adoption, and among the fifty most AI-cited domains only one had the file. Google's Gary Illyes confirmed Google does not support it and is not planning to. John Mueller compared it to the discredited keywords meta tag. Analysis of over 500 million AI bot events found the major crawlers overwhelmingly skip it and crawl the HTML directly.
Ship it anyway, it costs twenty minutes, but expect nothing from it. Your HTML is what actually gets read.
19. query, keyword, intent, category
Four words that get used interchangeably and should not be.
A query is what someone literally typed. is clay worth it for a 5 person team. Real queries are messy, long and specific.
A keyword is the shorthand for a group of similar queries. "clay pricing" as a keyword covers dozens of actual phrasings.
Intent is what the person wants to happen next.
Category is the topic cluster. "AI voice agents", "email deliverability". One category holds many keywords, each keyword holds many queries.
20. the four intents
| Intent | Example | What they want |
|---|---|---|
| Informational | "what is an AI SDR" | To understand |
| Commercial | "best AI SDR tools" | To compare before buying |
| Transactional | "apollo pricing" | To act now |
| Navigational | "clay login" | To reach a specific place |
The engine infers intent from the query and then only shows pages that match it. Search "best AI SDR tools" and you get comparison lists, not product homepages, because that query has been classified commercial, and a homepage does not satisfy it.
Your page can be excellent and still lose, simply for being the wrong type of page.
21. one page, one intent, one category
The rule that falls out of everything above.
A page trying to be both the explainer and the comparison satisfies neither, and the engine will pick someone else's page for both queries. Split them. Link them to each other. Each one then ranks for its own thing.
22. cannibalisation
Two of your own pages targeting the same query. They compete with each other, split the signals that would have gone to one page, and both rank worse than either would alone.
It is the most common self-inflicted wound on any site over about fifty pages, and it is invisible unless you keep a written record of which page owns which query.
23. SERP
Search Engine Results Page. The page you get after hitting enter.
Worth having a word for because that page is no longer ten blue links. One SERP might contain an AI Overview, ads, People Also Ask, a map pack, images, and then the organic results. Those components are SERP features, and several of them push the first organic result below the fold.
"Position 3" means less than it used to.
24. what decides rank
Simplified, but not wrong.
Relevance. does this page match the query and its intent. Mostly what is in the headings and text.
Authority. do other sites link to this one. A link is treated as a vote. Links from respected, topically related sites count for far more than volume.
Experience. does the page load fast, work on mobile, not ambush the reader. A tiebreaker, not a lever. It will not rescue a weak page.
Freshness. matters enormously for pricing pages and not at all for definitions.
25. E-E-A-T
Experience, Expertise, Authoritativeness, Trust. Google's framing for how its human quality raters judge a page. Not a score in the algorithm. A description of what the algorithm is trying to approximate.
Experience. the newest of the four. Have you actually done the thing? A review written by someone who used the product beats one assembled from spec sheets.
Expertise. do you know the subject.
Authoritativeness. does anyone else treat you as knowing it. This is the one you cannot self-declare.
Trust. the load-bearing one. Is the site honest about who runs it, how it makes money, and where its numbers came from.
What this cashes out to in practice is boring and effective: a real author with a real name, dates on everything, sources cited and linked, an about page, and disclosure of any commercial relationship. A directory that says plainly that it takes no vendor money is making an E-E-A-T argument.
26. SEO, AEO, GEO
Same underlying work, three different destinations.
SEO. get ranked in the links. Success = a position.
AEO, answer engine optimisation. Get pulled into the answer box, featured snippet or AI Overview. Needs clean question-and-answer structure and direct declarative sentences.
GEO, generative engine optimisation. Get cited inside ChatGPT, Claude or Perplexity answers. Success = your domain named as a source.
The practical difference: SEO rewards a page that ranks. AEO and GEO reward a page that is quotable. Specific numbers, clear comparisons, unambiguous claims that survive being lifted out of context.
Which means:
- Answer the question in the first sentence, then explain. Not the reverse.
- Use specific numbers rather than adjectives. "40,800 credits" survives being quoted; "great value" does not.
- Make each claim self-contained, so it still makes sense with the paragraph around it removed.
- Use real headings that match real questions.
One thing that has genuinely changed and is not yet widely understood: ranking and citation have come apart. Overlap between the top-10 organic results and AI Overview citations collapsed from around 75% in mid-2025 to somewhere between 17% and 38% by early 2026. Roughly 90% of pages cited by ChatGPT rank at position 21 or worse. Being uncitable is now a separate failure from ranking badly, and being unrankable no longer means being uncitable.
27. structure
Individual pages do not rank alone. They rank as part of a shape.
The shape is pillar and cluster. One broad page covers a topic at the level of the whole category. Around it sit narrower pages, each taking one specific question inside that topic. The pillar links down to each of them. Each of them links back up.
What that does: it tells the crawler these pages belong together, it passes authority around the group instead of letting it pool on one page, and it means a visitor who lands anywhere in the cluster can find the rest.
Internal linking is the mechanism, and it is the most underrated free lever there is. Rules that hold up:
- Link with descriptive text, not "click here". The words in the link tell the engine what the destination is about.
- Every page should be reachable within about three clicks of the homepage.
- New pages need links from existing pages, or nothing will find them for weeks.
- Link down from broad to narrow, and back up. Sideways between siblings only where a reader would genuinely want it.
Orphan pages, pages with no internal links pointing at them, are the failure mode. They exist, they are in the sitemap, and they are effectively invisible.
28. measured, observed, and modelled
The last one, and the one that will save you the most money.
Every number you will ever be shown about your site is one of three things.
Measured, first-party. Search Console and Bing Webmaster Tools. These are records of what actually happened. Real queries real people typed to reach your site, real impressions, real clicks. Free. This is ground truth.
Observed, third-party. SERP APIs. Someone ran the search and recorded what the page looked like at that moment. Real, but only a snapshot, and only of rankings.
Modelled, third-party. Every search volume figure in every SEO platform. Nobody outside Google knows how many people search a phrase per month. These are estimates built from clickstream panels and extrapolation, and vendors openly disagree with each other on the same keyword.
None of this makes modelled data useless. It makes it a different kind of number. Use it to choose between phrases for pages that do not exist yet, where you have no first-party data because you do not rank. Use first-party data for everything about pages you already have.
Confusing the three is how people spend $500 a month to be told something Search Console would have shown them for free.