<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
  <title>Nicasia Digital Solutions · Field notes</title>
  <link href="https://nicasiadigital.com/notes/feed.xml" rel="self"/>
  <link href="https://nicasiadigital.com/notes"/>
  <id>https://nicasiadigital.com/notes</id>
  <updated>2026-09-30T12:12:25+08:00</updated>
  <author><name>Nicasia Digital Solutions</name></author>
  <entry>
    <title>What tracking 146 homes on workers&#x27; phones taught us</title>
    <link href="https://nicasiadigital.com/notes/what-146-homes-taught-us-about-job-tracking"/>
    <id>https://nicasiadigital.com/notes/what-146-homes-taught-us-about-job-tracking</id>
    <published>2026-09-30T00:00:00+08:00</published>
    <updated>2026-09-30T12:12:25+08:00</updated>
    <summary>A list that offered 50 for a block of ten homes, a total inflated by 76 jobs that never existed, and old app screens calling a new server.</summary>
    <content type="html">&lt;p&gt;One housing project, 146 homes: 114 in the middle of their rows and 32 on the corners. Each home gets a set of metal items, a main gate, a side gate on the corners, railings, louvres on some house types, and every item passes through seven steps: measuring on site, ordering steel, fabrication, grinding and sealing, primer, installation and touch-up.&lt;/p&gt;
&lt;p&gt;The workers report each step from their phones, in Chinese, Bengali or Burmese. Their app runs inside Telegram, the records sit in a cloud database, and the owner follows every block and every home on a web dashboard without making a call. We built it for our own metal fabrication floor and it has been in daily use since.&lt;/p&gt;
&lt;p&gt;Most of what we learned in September 2026 came from four mistakes, all of them our own.&lt;/p&gt;
&lt;h2 id=&quot;a-block-of-ten-homes-should-offer-ten&quot;&gt;A block of ten homes should offer ten&lt;/h2&gt;
&lt;p&gt;On 4 September the owner opened the screen that assigns work in the factory. The quantity list offered 1, 2, 3, 4, 5, 6, 8, 10, 12, 20, 30 and 50. The block he had picked, B12, has ten homes. The item was louvres, which go on every home in that block, so the true maximum was ten. For the side gate, which only corner homes take, it was two.&lt;/p&gt;
&lt;p&gt;He asked why the list was not built from the database. He was right. Every number the list needed was already in it, and the screen that workers use to report finished jobs had been counting from it all along. The factory screen used a fixed list someone thought was close enough.&lt;/p&gt;
&lt;p&gt;A wrong list is worse than an ugly one, because people act on it. An order for 30 louvres from that screen would have sent the floor to cut 30 and the steel to be bought for 30.&lt;/p&gt;
&lt;p&gt;We changed three things that day:&lt;/p&gt;
&lt;ol&gt;&lt;li&gt;The block list shows only blocks where the item fits, with the number of homes beside each name: &amp;quot;B12 · 10&amp;quot;.&lt;/li&gt;&lt;li&gt;The quantity stops at what the records allow for that block and that item.&lt;/li&gt;&lt;li&gt;The server refuses the impossible order on its own. A list on a screen can be bypassed or out of date; the server is the last place that can say no.&lt;/li&gt;&lt;/ol&gt;
&lt;p&gt;We also split one message into two. &amp;quot;We cannot count this&amp;quot; and &amp;quot;this item is not counted per home&amp;quot; used to show the same thing on screen, a blank maximum. They are different facts, and the second one is normal: work for another company has no homes in our records. The screen now says which one it is.&lt;/p&gt;
&lt;h2 id=&quot;the-total-matters-as-much-as-the-count&quot;&gt;The total matters as much as the count&lt;/h2&gt;
&lt;p&gt;The same day, a check of every list against the database turned up a bigger number problem. The database said louvres go on all 146 homes. The bill of quantities, the contract&amp;#x27;s own count, said 70. The difference, 76, is exactly the number of type A homes, and louvres belong to type B only.&lt;/p&gt;
&lt;p&gt;That one field let workers report louvres on homes that never get them, and it put 76 jobs that did not exist into the project&amp;#x27;s total. When we corrected it, the project&amp;#x27;s progress moved from 4.2% to 4.7% overnight. Nobody worked faster. The total had shrunk to the truth.&lt;/p&gt;
&lt;p&gt;If a progress figure is going to reach a client, check the denominator against the bill of quantities before anyone reads the percentage.&lt;/p&gt;
&lt;h2 id=&quot;two-screens-deciding-the-same-thing-will-drift-apart&quot;&gt;Two screens deciding the same thing will drift apart&lt;/h2&gt;
&lt;p&gt;On 17 September a worker on another project finished one of the two homes in his block, reported it, and could not report the second. Tapping the block again to add the second home cleared his selection instead. The server had the right data the whole time. The trap was in the screen.&lt;/p&gt;
&lt;p&gt;While fixing it we checked every project for the same kind of trap, and found one the owner had not reported. The workers&amp;#x27; app and the owner&amp;#x27;s dashboard each decide which items a home gets, in separate code. After the louvres fix, the app judged from the house type recorded for each home. In the 146-home project the type was recorded for each block instead, so the app decided that none of the 70 type B homes took louvres, and nobody could report them. The dashboard looked fine. A hard-coded line there happened to give the right answer.&lt;/p&gt;
&lt;p&gt;The fix was one rule, used by both: when a block has no type recorded per home, every home takes the block&amp;#x27;s type. Before shipping it we listed what the dashboard decided for every home in the five projects tracked per home, 2,209 pairs of home and item, and compared the list before and after. The only judgments that moved were those 70 louvres, and no screen on the dashboard changed at all. That comparison is now a tool we keep.&lt;/p&gt;
&lt;h2 id=&quot;old-screens-keep-calling-the-new-server&quot;&gt;Old screens keep calling the new server&lt;/h2&gt;
&lt;p&gt;On 4 September the server&amp;#x27;s answer changed shape: the list of blocks went from plain names to name-and-count pairs, so the new screen could print &amp;quot;B12 · 10&amp;quot;. We shipped the new screen at the same time and tested both together.&lt;/p&gt;
&lt;p&gt;The owner&amp;#x27;s phone was not running the new screen. He had opened the app from an old Telegram message, and Telegram served him the page it had kept from before. That old page printed the new answer as text, and every block on his list read [object Object].&lt;/p&gt;
&lt;p&gt;Nothing crashed. That was the danger. Our tests always pair the newest screen with the newest server; people in the field do not. Telegram only changes the version a message opens when a new message is sent.&lt;/p&gt;
&lt;p&gt;The rule since then: a field the app already reads never changes shape. New information goes into a new field, and a new screen that cannot find it falls back to the old one and says what it does not know.&lt;/p&gt;
&lt;h2 id=&quot;what-to-ask-before-you-buy-a-site-tracking-system&quot;&gt;What to ask before you buy a site-tracking system&lt;/h2&gt;
&lt;ul&gt;&lt;li&gt;Where do its lists come from? Every choice a worker or supervisor taps should be counted from your project&amp;#x27;s records, per block and per home.&lt;/li&gt;&lt;li&gt;Does the server refuse what the records rule out, or only the screen?&lt;/li&gt;&lt;li&gt;Is the progress total built from your bill of quantities, and who checks it when the drawings and the bill disagree?&lt;/li&gt;&lt;li&gt;What happens when someone opens an old version of the app?&lt;/li&gt;&lt;li&gt;Can it tell &amp;quot;we do not know&amp;quot; from &amp;quot;this does not apply&amp;quot;?&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;We build these systems for other companies now, starting from their own records. &lt;a href=&quot;https://nicasiadigital.com/contact&quot;&gt;A private review&lt;/a&gt; looks at how the work moves in yours.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>Can ChatGPT and Claude read your website? How to check</title>
    <link href="https://nicasiadigital.com/notes/can-ai-assistants-read-your-website"/>
    <id>https://nicasiadigital.com/notes/can-ai-assistants-read-your-website</id>
    <published>2026-09-30T00:00:00+08:00</published>
    <updated>2026-09-30T12:01:25+08:00</updated>
    <summary>Our robots.txt welcomed every crawler, yet Cloudflare sent GPTBot a 403 and hid our email from AI crawlers. What each crawler does, and a one-line test.</summary>
    <content type="html">&lt;p&gt;On 19 September 2026 we tested one of our own websites the way an AI crawler meets it. Its robots.txt welcomed every crawler. Googlebot, Bingbot, OAI-SearchBot and PerplexityBot got the page. GPTBot, ClaudeBot and CCBot got 25 bytes of plain text from Cloudflare instead: &amp;quot;Your request was blocked.&amp;quot;&lt;/p&gt;
&lt;p&gt;On 30 September we found a quieter gap on this site, nicasiadigital.com. Every crawler could open the pages, but our email address reached them as &amp;quot;[email protected]&amp;quot;.&lt;/p&gt;
&lt;p&gt;We had switched on neither setting. Both came with Cloudflare, and neither shows when you open the site in a browser: a browser gets the page, and it runs the script that puts the address back.&lt;/p&gt;
&lt;h2 id=&quot;which-crawler-does-what&quot;&gt;Which crawler does what&lt;/h2&gt;
&lt;p&gt;The companies behind the AI assistants each run more than one crawler, and each crawler feeds something different. Blocking one can cost you a great deal or nothing, depending on which one it is. This is what each company&amp;#x27;s own documentation says, as of September 2026.&lt;/p&gt;
&lt;div class=&quot;tablewrap&quot;&gt;&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Crawler&lt;/th&gt;&lt;th&gt;Company&lt;/th&gt;&lt;th&gt;What it reads your site for&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;OAI-SearchBot&lt;/td&gt;&lt;td&gt;OpenAI&lt;/td&gt;&lt;td&gt;Showing your pages in ChatGPT search answers&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;GPTBot&lt;/td&gt;&lt;td&gt;OpenAI&lt;/td&gt;&lt;td&gt;Training future models&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;ChatGPT-User&lt;/td&gt;&lt;td&gt;OpenAI&lt;/td&gt;&lt;td&gt;Opening a page because a user asked ChatGPT to&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Claude-SearchBot&lt;/td&gt;&lt;td&gt;Anthropic&lt;/td&gt;&lt;td&gt;Improving Claude&amp;#x27;s search results&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;ClaudeBot&lt;/td&gt;&lt;td&gt;Anthropic&lt;/td&gt;&lt;td&gt;Training future models&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Claude-User&lt;/td&gt;&lt;td&gt;Anthropic&lt;/td&gt;&lt;td&gt;Opening a page because a user asked Claude to&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;PerplexityBot&lt;/td&gt;&lt;td&gt;Perplexity&lt;/td&gt;&lt;td&gt;Listing and linking sites in Perplexity&amp;#x27;s answers; Perplexity says it does not train on them&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Perplexity-User&lt;/td&gt;&lt;td&gt;Perplexity&lt;/td&gt;&lt;td&gt;Opening a page during an answer; it generally ignores robots.txt&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Google-Extended&lt;/td&gt;&lt;td&gt;Google&lt;/td&gt;&lt;td&gt;A name in robots.txt with no crawler behind it: it decides whether Gemini may train on your pages or use them to ground answers&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;CCBot&lt;/td&gt;&lt;td&gt;Common Crawl&lt;/td&gt;&lt;td&gt;An open archive of the web that AI companies have used to train models&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;Two points in that documentation matter most. OpenAI&amp;#x27;s page says a site that blocks OAI-SearchBot stops appearing in ChatGPT search answers, while blocking GPTBot only keeps its pages out of training. Google&amp;#x27;s says Google-Extended has no effect on Google Search.&lt;/p&gt;
&lt;p&gt;So on our site in September, ChatGPT search could still read us, because OAI-SearchBot got through. What we were losing was training: the next models would know less about the site.&lt;/p&gt;
&lt;h2 id=&quot;test-it-from-your-own-computer&quot;&gt;Test it from your own computer&lt;/h2&gt;
&lt;p&gt;This loop asks for your home page under each crawler&amp;#x27;s name and prints the status code. It runs in the Mac Terminal and on Linux; on Windows, use Git Bash. Replace the address with yours.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;for b in Googlebot bingbot GPTBot \
    OAI-SearchBot ChatGPT-User ClaudeBot \
    Claude-SearchBot Claude-User \
    PerplexityBot CCBot; do
  echo &amp;quot;$b $(curl -s -o /dev/null \
    -w &amp;#x27;%{http_code}&amp;#x27; -A \
    &amp;quot;Mozilla/5.0 (compatible; $b/1.0)&amp;quot; \
    https://www.example.com/)&amp;quot;
done&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;On 30 September at 11:05, Malaysia time, nicasiadigital.com answered 200 to all ten names. A 200 means the crawler gets the page. A 403 or a 503 means something in front of your site refuses it, whatever robots.txt says. If you would rather not open a terminal, &lt;a href=&quot;https://nicasiadigital.com/tools/ai-crawler-check&quot;&gt;our free crawler check&lt;/a&gt; runs the same test from a web page, along with the checks further down this note.&lt;/p&gt;
&lt;p&gt;The test has one limit. It sends each crawler&amp;#x27;s name from your computer, and some firewalls also check where a request comes from, so the real crawler can be treated differently. Take a 200 as a good sign. Take a 403 seriously: on our site in September it matched a setting that turns away the real crawlers too.&lt;/p&gt;
&lt;h2 id=&quot;where-cloudflare-keeps-the-switch&quot;&gt;Where Cloudflare keeps the switch&lt;/h2&gt;
&lt;p&gt;Since 1 July 2025, Cloudflare asks every new domain at sign-up whether to allow AI crawlers, and blocking is where it starts. Our domain moved to Cloudflare on 2 September 2026. Seventeen days later the training crawlers were still being turned away.&lt;/p&gt;
&lt;p&gt;In the dashboard, open the domain, then Security, then Settings, and find the Bot traffic group. &amp;quot;Configure AI bot policies&amp;quot; has three rows: Search, Agent and Training. Ours read Allow, Allow, Disallow. Training was the row sending 403 to GPTBot, ClaudeBot and CCBot. We set it to Allow at 15:10 on 19 September, and the next run of the loop showed 200 for all nine names we tested that day.&lt;/p&gt;
&lt;p&gt;At the bottom of the same panel, &amp;quot;Enable Bot Preference Sync&amp;quot; lets Cloudflare write its own lines into your robots.txt. We keep it off. The site has a robots.txt of its own, and two sets of rules would sooner or later disagree.&lt;/p&gt;
&lt;h2 id=&quot;the-email-address-crawlers-could-not-read&quot;&gt;The email address crawlers could not read&lt;/h2&gt;
&lt;p&gt;Cloudflare&amp;#x27;s Email Address Obfuscation was on for this domain, and we had not turned it on. It replaces each email address in a page with a link to &lt;code&gt;/cdn-cgi/l/email-protection&lt;/code&gt; and the words &amp;quot;[email protected]&amp;quot;. A small script then puts the real address back in the visitor&amp;#x27;s browser, so people never see the swap.&lt;/p&gt;
&lt;p&gt;A crawler that reads the HTML and stops there gets &amp;quot;[email protected]&amp;quot;. A December 2024 study by Vercel and MERJ found that GPTBot, ClaudeBot and PerplexityBot fetch pages without running their JavaScript, so for them the address is simply gone.&lt;/p&gt;
&lt;p&gt;We left the setting on and changed the pages. Cloudflare skips anything between &lt;code&gt;&amp;lt;!--email_off--&amp;gt;&lt;/code&gt; and &lt;code&gt;&amp;lt;!--/email_off--&amp;gt;&lt;/code&gt;, so our build now wraps every address in those two comments. Turning the feature off for the whole domain works too. Addresses inside &lt;code&gt;&amp;lt;script&amp;gt;&lt;/code&gt; tags and in the &lt;code&gt;&amp;lt;head&amp;gt;&lt;/code&gt; are never rewritten, which is why the address in our structured data was readable all along.&lt;/p&gt;
&lt;p&gt;To check yours, fetch the contact page the way a crawler does and count Cloudflare&amp;#x27;s markers:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;curl -s https://www.example.com/contact \
  -A &amp;quot;Mozilla/5.0 (compatible; GPTBot/1.0)&amp;quot; \
  | grep -c &amp;quot;email-protection&amp;quot;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Anything above 0 means at least one address on that page is hidden from crawlers. On our contact page it now prints 0.&lt;/p&gt;
&lt;h2 id=&quot;what-else-a-crawler-misses&quot;&gt;What else a crawler misses&lt;/h2&gt;
&lt;p&gt;The same study found that Gemini, through Googlebot, and AppleBot do render JavaScript. The OpenAI, Anthropic and Perplexity crawlers it measured did not. Anything your site adds after its scripts run is missing for them: prices loaded by a script, a menu built in the browser, reviews from a widget, sometimes the whole page. To see what they see, fetch the page with &lt;code&gt;curl&lt;/code&gt; and search the output for the words that matter to you.&lt;/p&gt;
&lt;h2 id=&quot;our-checklist-after-every-deploy&quot;&gt;Our checklist after every deploy&lt;/h2&gt;
&lt;ul&gt;&lt;li&gt;robots.txt allows the crawlers you want. Under &lt;code&gt;User-agent: *&lt;/code&gt; with &lt;code&gt;Allow: /&lt;/code&gt;, a crawler that is not named counts as allowed.&lt;/li&gt;&lt;li&gt;Every name in the loop above gets 200.&lt;/li&gt;&lt;li&gt;On Cloudflare, the AI bot policies say the same thing as robots.txt.&lt;/li&gt;&lt;li&gt;Your main words, prices and contact details are in the HTML before any script runs.&lt;/li&gt;&lt;li&gt;Your email address prints in the &lt;code&gt;curl&lt;/code&gt; check.&lt;/li&gt;&lt;li&gt;The sitemap lists every page, with dates that change only when a page&amp;#x27;s words change.&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;We run these on nicasiadigital.com after every deploy, from a script that lists anything that fails. If you want the same checks run on your site, &lt;a href=&quot;https://nicasiadigital.com/contact&quot;&gt;a private review&lt;/a&gt; covers them.&lt;/p&gt;
&lt;h2 id=&quot;sources&quot;&gt;Sources&lt;/h2&gt;
&lt;ul&gt;&lt;li&gt;OpenAI, &lt;a href=&quot;https://developers.openai.com/api/docs/bots&quot;&gt;Overview of OpenAI crawlers&lt;/a&gt;&lt;/li&gt;&lt;li&gt;Anthropic, &lt;a href=&quot;https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler&quot;&gt;Does Anthropic crawl data from the web, and how can site owners block the crawler?&lt;/a&gt;&lt;/li&gt;&lt;li&gt;Perplexity, &lt;a href=&quot;https://docs.perplexity.ai/guides/bots&quot;&gt;Perplexity crawlers&lt;/a&gt;&lt;/li&gt;&lt;li&gt;Google, &lt;a href=&quot;https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers&quot;&gt;Google&amp;#x27;s common crawlers&lt;/a&gt;&lt;/li&gt;&lt;li&gt;Cloudflare, &lt;a href=&quot;https://www.cloudflare.com/press/press-releases/2025/cloudflare-just-changed-how-ai-crawlers-scrape-the-internet-at-large/&quot;&gt;press release, 1 July 2025&lt;/a&gt;&lt;/li&gt;&lt;li&gt;Cloudflare, &lt;a href=&quot;https://developers.cloudflare.com/waf/tools/scrape-shield/email-address-obfuscation/&quot;&gt;Email Address Obfuscation&lt;/a&gt;&lt;/li&gt;&lt;li&gt;Vercel and MERJ, &lt;a href=&quot;https://vercel.com/blog/the-rise-of-the-ai-crawler&quot;&gt;The rise of the AI crawler&lt;/a&gt;, 17 December 2024&lt;/li&gt;&lt;/ul&gt;</content>
  </entry>
</feed>
