{"version":"https://jsonfeed.org/version/1.1","title":"Vic Demuzere - Web","home_page_url":"https://vic.demuzere.be/articles/tags/web/","feed_url":"https://vic.demuzere.be/articles/tags/web/feed.json","description":"Blog posts tagged Web by Vic Demuzere","language":"en","items":[{"id":"https://vic.demuzere.be/articles/bots-crawlers-and-geoblocks/","content_html":"<p>The last few years, web traffic from bots has increased rapidly.\nSome sources<sup class=\"footnote-ref\"><a href=\"#fn-1\" id=\"fnref-1\" data-footnote-ref>1</a></sup> claim that automated activity now accounts for over half of all web traffic.\nThis is even more pronounced on small websites like mine, as I don't receive many human visitors.</p>\n<p>The table below shows traffic on this website for the last few days:</p>\n<table>\n<thead>\n<tr>\n<th align=\"center\"></th>\n<th>Path</th>\n<th align=\"right\">Hits</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td align=\"center\">1</td>\n<td>/wp-login.php</td>\n<td align=\"right\">2751</td>\n</tr>\n<tr>\n<td align=\"center\">2</td>\n<td>/wp-admin/index.php</td>\n<td align=\"right\">2707</td>\n</tr>\n<tr>\n<td align=\"center\">3</td>\n<td>/wp-admin/edit.php</td>\n<td align=\"right\">2354</td>\n</tr>\n<tr>\n<td align=\"center\">4</td>\n<td>/wp-admin/profile.php</td>\n<td align=\"right\">2353</td>\n</tr>\n<tr>\n<td align=\"center\">5</td>\n<td>/wp-admin/plugins.php</td>\n<td align=\"right\">2351</td>\n</tr>\n<tr>\n<td align=\"center\">6</td>\n<td>/robots.txt</td>\n<td align=\"right\">1131</td>\n</tr>\n<tr>\n<td align=\"center\">7</td>\n<td>/</td>\n<td align=\"right\">994</td>\n</tr>\n<tr>\n<td align=\"center\">8</td>\n<td>/articles/feed.rss</td>\n<td align=\"right\">989</td>\n</tr>\n<tr>\n<td align=\"center\">9</td>\n<td>/articles/how-i-almost-lost-my-backups/</td>\n<td align=\"right\">314</td>\n</tr>\n<tr>\n<td align=\"center\">10</td>\n<td>/wp-content/themes/vic/style.css</td>\n<td align=\"right\">299</td>\n</tr>\n</tbody>\n</table>\n<p>This clearly shows two things:</p>\n<ul>\n<li>The first useful page for a real visitor is in sixth place<sup class=\"footnote-ref\"><a href=\"#fn-2\" id=\"fnref-2\" data-footnote-ref>2</a></sup>.</li>\n<li>My stylesheet has a low number of hits compared to the frontpage and my most recently published article.\nOne would expect each unique visitor to download the stylesheet at least once.</li>\n</ul>\n<p>That is to say, my own statistics show that I'm mainly serving our robot overlords.</p>\n<p>The situation is worse on other kinds of web apps, like my <a href=\"https://forgejo.org/\" title=\"Forgejo is a self-hosted lightweight software forge.\">Forgejo</a> instance.\nIt's a single-user instance, and most projects are private, but I do publish a bunch of open-source projects there.\nRecent bot traffic on applications like these is causing significant load on my server.</p>\n<h2 id=\"keeping-the-bots-out\">Keeping the bots out<a href=\"#keeping-the-bots-out\" aria-label=\"Link to heading 'Keeping the bots out'\" data-heading-content=\"Keeping the bots out\" class=\"anchor\"></a></h2>\n<p>Previously, the majority of crawlers were relatively decent.\nThey would crawl your website slowly, making sure not to cause too much load.\nTheir main purpose was serving search results to people looking for the content you publish.\nBringing down your website would only result in their own users landing on error pages.</p>\n<p>The recent <a href=\"https://en.wikipedia.org/wiki/Generative_AI\" title=\"Generative AI on Wikipedia\">Generative AI</a> (GenAI) craze changed this completely.\nMultiple companies are now racing to collect as much data as possible.\nThey no longer care what happens with the websites they're crawling.\nOnce they have collected enough content, their GenAI agents provide users with answers directly.\nThe source website is treated as irrelevant.\nObviously, as a website owner, I do not want these thankless robots on my property<sup class=\"footnote-ref\"><a href=\"#fn-3\" id=\"fnref-3\" data-footnote-ref>3</a></sup>!</p>\n<p>So I went through the standard ways of blocking them:</p>\n<ol>\n<li>Asking nicely in <a href=\"https://www.robotstxt.org/\" title=\"The Web Robots Pages\"><code>robots.txt</code></a>.</li>\n<li>Denying requests based on the <a href=\"https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Headers/User-Agent\" title=\"User-Agent HTTP header on Mozilla Developer Network\">User-Agent</a> header.</li>\n<li>Blocking IP ranges from known data centers and cloud providers.</li>\n</ol>\n<p>None of these methods work.</p>\n<p>GenAI companies frequently ignore <code>robots.txt</code> files.\nThey also usually disguise themselves as normal browsers instead of identifying as crawlers.\nMore recently, I've also noticed them increasingly using consumer IP space.\nCombined with rapidly changing addresses, this makes it impossible to block them by IP address alone.</p>\n<p>There are some promising new tools like <a href=\"https://anubis.techaro.lol/\" title=\"Anubis\">Anubis</a> that work<sup class=\"footnote-ref\"><a href=\"#fn-4\" id=\"fnref-4\" data-footnote-ref>4</a></sup>, for now.</p>\n<h2 id=\"giving-up\">Giving up<a href=\"#giving-up\" aria-label=\"Link to heading 'Giving up'\" data-heading-content=\"Giving up\" class=\"anchor\"></a></h2>\n<p>Making sure websites stay online is part of my job.\nI don't want to have to go through the same hassle for my own web apps when I get home.\nI'm tired, and the few things I'm hosting just aren't worth it.</p>\n<p>So, I decided to take two different approaches:</p>\n<ol>\n<li>\n<p>Websites that are <strong>meant to be public</strong>, like the one you're reading now.\nI've always optimized these for high traffic, and will continue to do so.\nThese are built to withstand much more than the AI crawlers I'm currently seeing.\n<strong>I'm no longer blocking any bots on public websites.</strong></p>\n</li>\n<li>\n<p>Websites (and web apps) that are <strong>primarily meant for myself</strong>.\nThese are <strong>now geoblocked</strong> and only accept traffic from Belgium.\nAs these are mainly open source applications that I don't maintain myself, they are harder to optimize.\nThis includes my <a href=\"https://forgejo.org/\" title=\"Forgejo is a self-hosted lightweight software forge.\">Forgejo</a> instance.\nIt's sad that part of my open source code disappears behind a firewall.\nI hope to open it up again once AI companies start behaving more responsibly.</p>\n</li>\n</ol>\n<p>My geoblock page contains contact information.\nSo if people <em>really want to go through the hassle</em> of requesting access (which I very much doubt), they can be allowed in.</p>\n<section class=\"footnotes\" data-footnotes>\n<ol>\n<li id=\"fn-1\">\n<p>Every article I've found claiming this reads like marketing material for a web security firm.\nFor that reason, I decided not to link to any of them.\nBased on my own traffic stats though, I don't doubt their numbers. <a href=\"#fnref-1\" class=\"footnote-backref\" data-footnote-backref data-footnote-backref-idx=\"1\" aria-label=\"Back to reference 1\">↩</a></p>\n</li>\n<li id=\"fn-2\">\n<p>I do read robots.txt on some sites I visit.\nIt's a treasure trove of information about the website you're visiting.\nBut that's probably not normal human behaviour. <a href=\"#fnref-2\" class=\"footnote-backref\" data-footnote-backref data-footnote-backref-idx=\"2\" aria-label=\"Back to reference 2\">↩</a></p>\n</li>\n<li id=\"fn-3\">\n<p>I don't really mind AI crawlers, as long as they behave.\nThey usually don't. <a href=\"#fnref-3\" class=\"footnote-backref\" data-footnote-backref data-footnote-backref-idx=\"3\" aria-label=\"Back to reference 3\">↩</a></p>\n</li>\n<li id=\"fn-4\">\n<p>Anubis is great!\nEspecially since it lets through the bots that clearly identify as a bot. <a href=\"#fnref-4\" class=\"footnote-backref\" data-footnote-backref data-footnote-backref-idx=\"4\" aria-label=\"Back to reference 4\">↩</a></p>\n</li>\n</ol>\n</section>\n<div class=\"footer\"><hr><p>Thanks for subscribing to my feed!</p><p>Feel free to e-mail comments or corrections to <a href=\"mailto:vic@demuzere.be?subject=RE%3A%20Bots%20are%20dominating%20the%20web\">vic@demuzere.be</a>.</p></div>","url":"https://vic.demuzere.be/articles/bots-crawlers-and-geoblocks/","date_published":"2026-05-14T13:41:51Z","date_modified":"2026-05-14T13:41:51Z","language":"en"},{"id":"https://vic.demuzere.be/articles/matching-public-private-key/","content_html":"<p>Customers regularly send me x509 key pairs at work<sup class=\"footnote-ref\"><a href=\"#fn-1\" id=\"fnref-1\" data-footnote-ref>1</a></sup>.\nMost of the time these have been forwarded a few times already, and people tend to mess up the files.\nThis means that the public and private keys don't match anymore, and I have to ask for a new pair.</p>\n<p>Luckily, we can use OpenSSL to check whether a public and private key match.</p>\n<h2 id=\"for-rsa-keys\">For RSA keys<a href=\"#for-rsa-keys\" aria-label=\"Link to heading 'For RSA keys'\" data-heading-content=\"For RSA keys\" class=\"anchor\"></a></h2>\n<p>The classic way to compare RSA keys is to check whether the <strong>modulus</strong> is the same, as this is the only part that is shared between the public and private keys.\nIf the modulus isn't the same, the keys can't form a pair.</p>\n<p>Read the modulus from the <strong>public key</strong>:</p>\n<pre><code>openssl rsa -pubin -in public.pem -modulus -noout | openssl sha1\n</code></pre>\n<p>Then do the same with the <strong>private RSA</strong> key:</p>\n<pre><code>openssl rsa -in private.pem -modulus -noout | openssl sha1\n</code></pre>\n<p>If the two values match, the key files form a pair.</p>\n<p>Note that we're piping the actual output of the <code>-modulus</code> command to <code>openssl sha1</code> to get a shorter hash of the modulus.\nThis makes it easier to manually compare the two values, but it's not strictly necessary.\nYou could also compare the output directly if you're feeling brave.</p>\n<h2 id=\"for-ec-keys\">For EC keys<a href=\"#for-ec-keys\" aria-label=\"Link to heading 'For EC keys'\" data-heading-content=\"For EC keys\" class=\"anchor\"></a></h2>\n<p>Elliptic curve keys don't have an easy-to-compare modulus like the RSA keys above had.\nHowever, it's always possible to generate the public key<sup class=\"footnote-ref\"><a href=\"#fn-2\" id=\"fnref-2\" data-footnote-ref>2</a></sup> part when we have the private key.</p>\n<p>Using the <strong>private key</strong>, generate (a hash of) the public key:</p>\n<pre><code>openssl pkey -pubout -in private.pem | openssl sha1\n</code></pre>\n<p>Then compare the output with the <strong>public key</strong> file you have.\nWe need to make <code>openssl</code> load it and print it out again to make sure we get the exact same format.</p>\n<pre><code>openssl x509 -pubkey -in public.pem -noout | openssl sha1\n</code></pre>\n<p>If the two values match, these keys form a pair.</p>\n<section class=\"footnotes\" data-footnotes>\n<ol>\n<li id=\"fn-1\">\n<p>Please don't do this. Ask the company that needs a certificate to generate a private key and provide you with a <a href=\"https://en.wikipedia.org/wiki/Certificate_signing_request\">Certificate Signing Request (CSR)</a>. This way, you don't have to handle the private key at all. <a href=\"#fnref-1\" class=\"footnote-backref\" data-footnote-backref data-footnote-backref-idx=\"1\" aria-label=\"Back to reference 1\">↩</a></p>\n</li>\n<li id=\"fn-2\">\n<p>When people say &quot;public key&quot; in the context of a x509 pair they usually refer to the certificate. The public key we're extracting here is the actual cryptographic public key, it doesn't contain the extra data like domain names that a certificate usually holds. You can't recover the full certificate when you only have the private key. <a href=\"#fnref-2\" class=\"footnote-backref\" data-footnote-backref data-footnote-backref-idx=\"2\" aria-label=\"Back to reference 2\">↩</a></p>\n</li>\n</ol>\n</section>\n<div class=\"footer\"><hr><p>Thanks for subscribing to my feed!</p><p>Feel free to e-mail comments or corrections to <a href=\"mailto:vic@demuzere.be?subject=RE%3A%20Check%20whether%20public%20and%20private%20key%20match\">vic@demuzere.be</a>.</p></div>","url":"https://vic.demuzere.be/articles/matching-public-private-key/","date_published":"2025-02-06T21:00:00Z","date_modified":"2025-02-07T10:59:22Z","language":"en"},{"id":"https://vic.demuzere.be/articles/hiding-outdated-articles/","content_html":"<p><a href=\"https://kevquirk.com\" title=\"Kev Quirk\">Kev Quirk</a> recently wrote about <a href=\"https://kevquirk.com/blog/on-removing-content\" title=\"On Removing Content; Kev Quirk\">removing content from public blogs</a>:</p>\n<blockquote>\n<p>Generally speaking, I don't delete content from this site.\nHaving said that, if I posted something that I later feel is particularly egregious, I think I probably would.\nI personally don't think that a website should be a permanent record - how can it be?\nNothing last forever.</p>\n</blockquote>\n<p>This got me thinking about what to do with some of my older articles.\nI believe that URLs should be stable.\nThe internet is a large place, there’s no way to know who might be linking to any of my webpages.\nRemoving a page may also <em>break other websites</em>: someone may have written about it.\nDeleting the page would remove the context of their reference.</p>\n<p>However, I’ve been working on this website in some form or another since 2007.\nTechnical articles tend to get outdated, and people change their minds about things they’ve written.\nSome of the articles on this website no longer reflect what I stand for.</p>\n<p>Articles older than two years already include a banner noting their age, but that’s not always enough.\nNowadays, people often form an opinion based on the title alone — especially on social media.\nIt’s unlikely they’d notice a note about how old the article is.</p>\n<p><strong>So I decided to <em>hide</em> certain articles.</strong>\nHidden articles are no longer linked on my website and don't appear in RSS feeds.\nYou can’t find them by clicking around.\nBut if you already have the link, the content is still accessible.\nBecause, as they say, <a href=\"https://www.w3.org/Provider/Style/URI\" title=\"Cool URIs don't change\">cool URIs don't change</a>.</p>\n<div class=\"footer\"><hr><p>Thanks for subscribing to my feed!</p><p>Feel free to e-mail comments or corrections to <a href=\"mailto:vic@demuzere.be?subject=RE%3A%20Cool%20URIs%20don%27t%20change%2C%20but%20don%27t%20need%20spotlights\">vic@demuzere.be</a>.</p></div>","url":"https://vic.demuzere.be/articles/hiding-outdated-articles/","date_published":"2024-11-10T10:04:30Z","date_modified":"2024-11-11T18:18:07Z","language":"en"},{"id":"https://vic.demuzere.be/articles/not-so-redesign/","content_html":"<p>Over the past few months I've been working on a <strong>new stylesheet</strong> for this website, and today I finally pushed the redesign branch to production!\nThe previous CSS code was largely written <em>almost 10 years ago</em> and only received small updates and enhancements.\nSince then, new features like <a href=\"https://developer.mozilla.org/en-US/docs/Glossary/Flexbox\" title=\"MDN Glossary: Flexbox\">flexbox</a> and <a href=\"https://developer.mozilla.org/en-US/docs/Glossary/Grid\" title=\"MDN Glossary: Grid\">grid</a>s were added and made creating responsive websites much easier.</p>\n<p>Turns out, the new design looks <strong>exactly the same as the old one</strong>!</p>\n<p>A few things did change though:</p>\n<ul>\n<li>Wrote plain CSS instead of <a href=\"http://lesscss.org/\" title=\"LESS: CSS, with just a little more\">LESS</a>.</li>\n<li>Stopped using <code>float</code> to position elements in favor of <a href=\"https://developer.mozilla.org/en-US/docs/Glossary/Flexbox\" title=\"MDN Glossary: Flexbox\">flexbox</a>.</li>\n<li>Followed the <a href=\"https://www.w3.org/WAI/standards-guidelines/wcag/\" title=\"Web Content Accessibility Guidelines (WCAG)\">Web Content Accessibility Guidelines</a> (WCAG).</li>\n<li>Made the stylesheet a lot smaller.</li>\n<li>Focused on readability throughout the website.</li>\n</ul>\n<p>I also tried to get some insight in how people used the website:</p>\n<h2 id=\"a-year-in-statistics\">A year in statistics<a href=\"#a-year-in-statistics\" aria-label=\"Link to heading 'A year in statistics'\" data-heading-content=\"A year in statistics\" class=\"anchor\"></a></h2>\n<p>Most of the traffic to this website is routed through <a href=\"https://www.cloudflare.com/cdn/\" title=\"Cloudflare CDN\">Cloudflare's CDN</a>.\nI've configured it to cache everything, which means the access logs on my web server record only part of the actual traffic.\nThis creates an interesting dataset: I can see which pages are most visited, but I can't see spikes due to pages being shared somewhere.\nFor the same reason, my statistics about browsers, operating systems, referring websites and returning visitors are completely skewed.</p>\n<h3 id=\"popular-pages\">Popular pages<a href=\"#popular-pages\" aria-label=\"Link to heading 'Popular pages'\" data-heading-content=\"Popular pages\" class=\"anchor\"></a></h3>\n<p>Articles are the most popular pages on this website.\nThat makes sense, as those tend to show up in searches and are shared more frequently than others.\nMy most popular articles are technical ones:</p>\n<ol>\n<li><a href=\"https://vic.demuzere.be/articles/using-systemd-user-units/\">Managing services for non-root users with systemd</a></li>\n<li><a href=\"https://vic.demuzere.be/articles/curriculum-vitae-cv-with-latex-moderncv/\">Creating a C.V. with LaTeX and moderncv</a></li>\n<li><a href=\"https://vic.demuzere.be/articles/golang-makefile-crosscompile/\">Cross compiling Go applications with Make</a></li>\n</ol>\n<p>That means I need to write more articles!</p>\n<h3 id=\"popular-404-errors\">Popular 404 errors<a href=\"#popular-404-errors\" aria-label=\"Link to heading 'Popular 404 errors'\" data-heading-content=\"Popular 404 errors\" class=\"anchor\"></a></h3>\n<p>The most common 404s appear to be scanners looking for vulnerabilities.\nWordPress related pages like <code>wp-login.php</code>, <code>xmlrpc.php</code> and several files inside the <code>wp-includes</code> directory lead the race, followed by Joomla's <code>administrator</code> directory.</p>\n<p>My work to fix broken links (using data from the access logs and Google Search Console) seems to have paid off.\nI've also been trying not to delete or move pages for a few years now.\nWhen a page is no longer relevant, I replace it with a small message pointing people to alternatives, like I did with my <a href=\"https://vic.demuzere.be/otr/\" title=\"My off-the-record fingerprints\">OTR fingerprint listing</a>.</p>\n<p>Another interesting 404 page is <code>/irc:irc.quakenet.org/sorcix,isnick</code> which seems to be caused by a spider that doesn't understand <code>irc:</code> URLs.</p>\n<h2 id=\"future-changes\">Future changes<a href=\"#future-changes\" aria-label=\"Link to heading 'Future changes'\" data-heading-content=\"Future changes\" class=\"anchor\"></a></h2>\n<p>I've been looking into <a href=\"https://indieweb.org\" title=\"The IndieWeb\">IndieWeb</a> after reading a few blog posts by <a href=\"https://jlelse.blog\" title=\"Jan-Lukas Else's weblog\">Jan-Lukas Else</a>.\nPeople can reply to your articles by linking back from their own website, or by using one of the provided IndieWeb-compatible services.\nAs I'm not active on mainstream social media, this may be a great way to connect with like-minded people.</p>\n<p>Features I want to roll out next:</p>\n<ul>\n<li>RSS feed for articles</li>\n<li>Overview pages for every tag, making it easier to find related articles</li>\n<li>Support for Webmentions</li>\n</ul>\n<div class=\"footer\"><hr><p>Thanks for subscribing to my feed!</p><p>Feel free to e-mail comments or corrections to <a href=\"mailto:vic@demuzere.be?subject=RE%3A%20The%20mayor%20CSS%20not-so-redesign\">vic@demuzere.be</a>.</p></div>","url":"https://vic.demuzere.be/articles/not-so-redesign/","date_published":"2020-06-14T00:12:27Z","date_modified":"2024-11-11T18:24:32Z","language":"en"},{"id":"https://vic.demuzere.be/articles/printer-friendly-website/","content_html":"<p>Certain visitors of your blog may want to print your articles.\nBy default, most browsers will present a <strong>simplified version</strong> of your website for printing.\nBackground images may be hidden, saving a lot of ink.\nBut we can do better: usually people just want to <strong>print your content</strong>, not the entire website.</p>\n<p>Printing things is not that common,\nso it's really not worth it to spend much time worrying about how your website looks on paper.\nThere are a few simple things that make a lot of difference, though!</p>\n<h2 id=\"style-for-printing\">Style for printing<a href=\"#style-for-printing\" aria-label=\"Link to heading 'Style for printing'\" data-heading-content=\"Style for printing\" class=\"anchor\"></a></h2>\n<p>Modern browsers <a href=\"http://caniuse.com/#feat=css-queries\" title=\"Media Queries at CanIUse.com\">support media queries</a>,\na way to limit part of your stylesheet to devices that match certain criteria.\nWhile this feature is most used to scale your website to the size of multiple devices,\nwe can also use it to define <strong>print styles</strong>.</p>\n<pre><code>@media print {\n\t/* insert CSS for printing here */\n}\n</code></pre>\n<p>We want these styles at the very <strong>bottom of our stylesheet</strong>:\nThey have to override everything that we have defined before.</p>\n<h3 id=\"older-browsers\">Older browsers<a href=\"#older-browsers\" aria-label=\"Link to heading 'Older browsers'\" data-heading-content=\"Older browsers\" class=\"anchor\"></a></h3>\n<p>In case you want to support Internet Explorer 8\nand lower (or prefer your print styles in a separate file),\nthe alternative is using a <code>link</code> tag with <code>media</code> attribute:</p>\n<pre><code>&lt;link rel=&quot;stylesheet&quot; href=&quot;print-all-the-things.css&quot; media=&quot;print&quot; /&gt;\n</code></pre>\n<p>When using a separate stylesheet, you no longer have to use the media queries in the examples below.</p>\n<h2 id=\"hiding-unnecessary-parts-of-the-page\">Hiding unnecessary parts of the page<a href=\"#hiding-unnecessary-parts-of-the-page\" aria-label=\"Link to heading 'Hiding unnecessary parts of the page'\" data-heading-content=\"Hiding unnecessary parts of the page\" class=\"anchor\"></a></h2>\n<p>Usually you'll want to get rid of everything except the content of your website.\nNo header, footer, sidebar, sharing widgets, advertisements..\nMost of this should be possible <em>by changing your stylesheet only</em>.\nYou might want to introduce a few new classes to manage visibility, that'll come in handy!</p>\n<pre><code>.print {\n\tdisplay: none;\n}\n@media print {\n\t#header, #footer, #sidebar, .share, .ads, .no-print {\n\t\tdisplay: none;\n\t}\n\t.print {\n\t\tdisplay: block;\n\t}\n}\n</code></pre>\n<h2 id=\"page-margins\">Page margins<a href=\"#page-margins\" aria-label=\"Link to heading 'Page margins'\" data-heading-content=\"Page margins\" class=\"anchor\"></a></h2>\n<p>Another really simple tweak is the page margin.\nI prefer margins <em>a bit wider than the defaults</em> that Chrome gave me.\nNote that most browsers add some information in the page header and footer;\nMaking the top and bottom margins a bit bigger helps separating your content from the page info.</p>\n<pre><code>@media print {\n\t@page {\n\t\tmargin: 2cm 2cm 2cm 1.5cm;\n\t}\n}\n</code></pre>\n<div class=\"footer\"><hr><p>Thanks for subscribing to my feed!</p><p>Feel free to e-mail comments or corrections to <a href=\"mailto:vic@demuzere.be?subject=RE%3A%20Making%20your%20blog%20printer-friendly\">vic@demuzere.be</a>.</p></div>","url":"https://vic.demuzere.be/articles/printer-friendly-website/","date_published":"2016-06-18T14:34:00Z","date_modified":"2025-03-09T10:03:36Z","language":"en"},{"id":"https://vic.demuzere.be/articles/caching-static-website-with-cloudflare/","content_html":"<p><a href=\"https://www.cloudflare.com\" title=\"Cloudflare\">Cloudflare</a> is a cloud-based solution to protect your website against attacks.\nOne way to mitigate (DDoS) attacks is providing a geographically distributed network of <a href=\"https://en.wikipedia.org/wiki/Point_of_presence\" title=\"Point of Presence\">POP</a>s with a lot of bandwidth.\nAnd that's exactly what they did: Cloudflare maintains a <a href=\"https://en.wikipedia.org/wiki/Content_delivery_network\" title=\"Content Delivery Network (CDN)\">content delivery network</a> with <a href=\"https://www.cloudflare.com/network-map\" title=\"Cloudflare's Network\">30 data centers</a> around the world.\nEven if you have no reason to suspect being attacked any time soon, the speed improvement is well worth it.</p>\n<p>Setting up their service is pretty straightforward. Move your domain to their nameservers and select the subdomains that have to be protected.\nImportant to note: Using the CDN routes all traffic to your website through the Cloudflare network. This means that you're basically allowing them\nto do a <a href=\"https://en.wikipedia.org/wiki/Man-in-the-middle_attack\" title=\"Man-in-the-middle Attack\">man-in-the-middle attack</a> on your visitors. This article is aimed at static websites, so it doesn't really matter, but it's something to keep in mind.</p>\n<h2 id=\"caching\">Caching<a href=\"#caching\" aria-label=\"Link to heading 'Caching'\" data-heading-content=\"Caching\" class=\"anchor\"></a></h2>\n<p>By default, <a href=\"https://www.cloudflare.com\" title=\"Cloudflare\">Cloudflare</a> caches resources that are supposed to be static on most websites. This includes images, stylesheets and javascripts.\nYour server is contacted for every pageview, but most assets are offloaded to their edge servers.\nThis is great for a dynamic website, as we want to make sure that visitors see their personalized version of the website.\nStatic websites, however, may benefit from being cached completely.</p>\n<p>First things first. In the Cloudflare site settings page, we see three cache levels:</p>\n<ul>\n<li>\n<p><strong>Aggressive</strong>: The complete URL is used to identify a resource.</p>\n<p>http://example.com/image.png?including=query-string</p>\n</li>\n<li>\n<p><strong>Simplified</strong>: The query string is ignored when identifying a resource.</p>\n<p>http://example.com/image.png?ignore=this</p>\n</li>\n<li>\n<p><strong>Basic</strong>: Only caches resources without query string.</p>\n<p>http://example.com/image.png</p>\n</li>\n</ul>\n<p>The cache level setting only applies to those resources that Cloudflare decided to cache.\nI usually set this to <em>Aggressive</em> for static websites, as it allows me to change a hash in the query string to invalidate the cache when updating a page.</p>\n<p>A second interesting setting is the <em>minimum expire TTL</em>, this allows you to overwrite the expire time in the <code>cache-control</code> header.\nWhen setting up a static website, I expect these headers to be set correctly on the origin server, but this may come in handy when you have little control.</p>\n<h2 id=\"cache-everything\">Cache Everything<a href=\"#cache-everything\" aria-label=\"Link to heading 'Cache Everything'\" data-heading-content=\"Cache Everything\" class=\"anchor\"></a></h2>\n<p>The setting we're looking for is hidden deeper in the Cloudflare admin panel. In order to make their CDN cache everything, you need to create a page rule.\nPage rules define custom settings for a given subset of your website. Click the gear icon next to your website and select <em>Page Rules</em>. Between the rule settings an item named <em>Custom caching</em> appears. Set this to <em>Cache everything</em> to, well, cache everything. This only applies to the URL pattern specified for the page rule.</p>\n<p>Two other interesting settings:</p>\n<ul>\n<li>\n<p><strong>Edge cache expire TTL</strong>: Specifies how long a resource has to be cached on the Cloudflare edge servers. The longer they are allowed to cache it, the more\nbandwidth you'll save. Decide how big this value may be depending on how much you update your static website. You may always clear their edge cache using the\nadministration panel. (This setting is only available when <em>Cache everything</em> is selected.)</p>\n</li>\n<li>\n<p><strong>Browser cache expire TTL</strong>: This setting overwrites the <em>minimum expire TTL</em> from the global site settings, for this specific URL pattern.</p>\n</li>\n</ul>\n<p>Try it out! Depending on the number of visitors, caching everything may save quite a lot of bandwidth! If you have only a few visitors per month, you may not\neven see it, especially when they are geographically distributed. It looks like Cloudflare does not share resources between edge servers. If a resource\nis not available in the <a href=\"https://en.wikipedia.org/wiki/Point_of_presence\" title=\"Point of Presence\">POP</a> closest to your visitor, it contacts the your server.</p>\n<div class=\"footer\"><hr><p>Thanks for subscribing to my feed!</p><p>Feel free to e-mail comments or corrections to <a href=\"mailto:vic@demuzere.be?subject=RE%3A%20Caching%20a%20static%20website%20on%20Cloudflare\">vic@demuzere.be</a>.</p></div>","url":"https://vic.demuzere.be/articles/caching-static-website-with-cloudflare/","date_published":"2015-01-17T23:50:00Z","date_modified":"2026-05-13T18:17:26Z","language":"en"},{"id":"https://vic.demuzere.be/articles/trackerless-social-sharing/","content_html":"<p>We seem to live in the golden age of social media: Sharing every useless article you've read with the world is important. Most visitors are of the lazy kind, though, so we need to make this as easy as possible. Fortunately, social media sites provide us with widgets to do exactly that! Phew, we're done here, right? ..right?</p>\n<p>The problem with these widgets is that we also allow social media companies to follow our visitors. The primary source of revenue for these companies is advertising. Knowing which sites their members visit is a valuable asset. <strong>But we don't want evil corporations to spy on our visitors, do we?</strong> The solution for this problem was invented in <em>1993</em>, and is called a <a href=\"https://developer.mozilla.org/en-US/docs/Web/HTML/Element/a\" title=\"Hyperlink, defined by the Mozilla Developer Network (MDN)\">hyperlink</a>. It was designed to send people from one page to the next.</p>\n<p>Not convinced yet? Well, this probably means you don't value your visitors enough. But there is another reason: <strong>Social sharing widgets slow down your website</strong>. Every widget has to be loaded from an external website, causing a new HTTP request. Hyperlinks just sit there, waiting to be clicked!</p>\n<p>Last but not least, social sharing hyperlinks can be styled as you wish. Widgets are a pain in the ass to integrate in your website. Most websites simply have them all together in a box. Choose your own social media icon for your links, or use my <a href=\"https://github.com/sorcix/socialvectors\" title=\"Vector icons for social media sites\">Social media vectors</a>.</p>\n<h2 id=\"so-where-do-you-need-to-link\">So, where do you need to link?<a href=\"#so-where-do-you-need-to-link\" aria-label=\"Link to heading 'So, where do you need to link?'\" data-heading-content=\"So, where do you need to link?\" class=\"anchor\"></a></h2>\n<ul>\n<li>\n<p>Twitter</p>\n<p><code>https://twitter.com/intent/tweet?url=URL</code></p>\n</li>\n<li>\n<p>Facebook:</p>\n<p><code>https://facebook.com/sharer/sharer.php?u=URL</code></p>\n</li>\n<li>\n<p>Google+:</p>\n<p><code>https://plus.google.com/share?url=URL</code></p>\n</li>\n<li>\n<p>Linkedin:</p>\n<p><code>https://www.linkedin.com/shareArticle?mini=true&amp;url=URL&amp;title=TITLE</code></p>\n</li>\n<li>\n<p>Reddit:</p>\n<p><code>https://www.reddit.com/submit?url=URL</code></p>\n</li>\n</ul>\n<div class=\"footer\"><hr><p>Thanks for subscribing to my feed!</p><p>Feel free to e-mail comments or corrections to <a href=\"mailto:vic@demuzere.be?subject=RE%3A%20Social%20sharing%20buttons%20without%20the%20trackers\">vic@demuzere.be</a>.</p></div>","url":"https://vic.demuzere.be/articles/trackerless-social-sharing/","date_published":"2014-04-21T18:00:00Z","date_modified":"2023-11-28T19:55:32Z","language":"en"}]}