Skip to content

Update Deno documentation scraper - #2681

Open
srpatcha wants to merge 3 commits into
freeCodeCamp:mainfrom
srpatcha:chore/add-gitattributes
Open

Update Deno documentation scraper#2681
srpatcha wants to merge 3 commits into
freeCodeCamp:mainfrom
srpatcha:chore/add-gitattributes

Conversation

@srpatcha

Copy link
Copy Markdown

Changes

chore: add .gitattributes for line ending normalization

Ensure consistent line endings and proper diff handling
for text and binary files.

Signed with GPG.

Ensure consistent line endings and proper diff handling
for text and binary files.
@srpatcha
srpatcha requested a review from a team as a code owner April 25, 2026 02:42
Add a new UrlScraper for Deno standard library and runtime documentation
(lib/docs/scrapers/deno.rb) with:

- Deno docs site scraping (docs.deno.com)
- Page parsing with Nokogiri (main/article content extraction)
- Link resolution (relative to absolute URL conversion)
- Version handling with semver normalization (v2 and v1 support)
- Module categorization (Web APIs, I/O, File System, Network, etc.)
- Code example extraction with language detection
- HTML filter pipeline (clean_html and entries filters)

Also includes:
- Minitest test class for scraper configuration validation
- Bug fix: replace File.open(path).read with File.read(path) in
  sprites.thor to prevent unclosed file handle leak

Signed-off-by: Srikanth Patchava <spatchava@meta.com>
@simon04 simon04 changed the title chore: add .gitattributes for line ending normalization Update Deno documentation scraper May 26, 2026
@simon04

simon04 commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

The scraper needs fine tuning:

image image

@srpatcha

Copy link
Copy Markdown
Author

Thanks for the review and the screenshots @simon04! You're right that the output needs refinement.

From the screenshots I can see a few issues to address:

  1. Residual navigation/UI chrome that isn't being stripped by the current clean_html selectors
  2. Code block language detection could be more accurate

I'll iterate on the clean_html.rb and entries.rb filters to clean up the rendering. Could you confirm which Deno docs section/URL the screenshots were captured from so I can target the exact page structure when tuning the selectors? That will help me verify the fixes against the same pages you're seeing.

@simon04 simon04 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

.

Remove source-site navigation and controls that leaked into generated API and runtime pages, while preserving the source language metadata DevDocs uses for syntax highlighting. Add representative regression fixtures for the reviewer-reported page shapes.

Test Plan:
- Ran the focused Deno suite: 17 runs, 73 assertions, 0 failures.
- Ran the remaining Ruby suite excluding two independently failing Windows-only baseline files: 665 runs, 927 assertions, 0 failures.
- Compared old and new filters on current Network and @std/fmt source pages; all measured chrome counts dropped to zero while TypeScript, JavaScript, and shell language metadata was preserved.
- Ran Ruby syntax checks and git diff --check.
@srpatcha

srpatcha commented Sep 8, 2026

Copy link
Copy Markdown
Author

@simon04 I pushed d17aafb to address the requested scraper fine-tuning without waiting for the missing source URL. I matched the screenshots to the current pages and validated the previous vs updated filter on both:

  • https://docs.deno.com/api/deno/network/: breadcrumbs 1 -> 0, symbol-kind badges 64 -> 0, code-copy buttons 23 -> 0, feedback sections 1 -> 0; all 23 TypeScript blocks still emit pre[data-language="ts"].
  • https://docs.deno.com/runtime/reference/std/fmt/: copy-page controls 1 -> 0, code-copy buttons 2 -> 0, mobile “On this page” blocks 1 -> 0, heading anchors 1 -> 0, feedback sections 1 -> 0; JavaScript and shell blocks remain js and sh.

The filter now scopes output to main#content article, removes the confirmed source-site chrome, and transfers each live language-* class to the enclosing pre[data-language] instead of defaulting every block to TypeScript. I also added reduced API/runtime fixtures and filter/entry regressions.

Validation:

  • Focused Deno tests: 17 runs, 73 assertions, 0 failures/errors.
  • Remaining Ruby suite excluding two independently reproduced Windows-only baseline files: 665 runs, 927 assertions, 0 failures/errors.
  • Full suite reached 761 runs; only the existing Windows case-sensitivity URL assertion and FileStore test/tmp path errors failed when run independently.
  • Ruby syntax checks, editor diagnostics, change validation, and git diff --check passed.
  • The live Thor scraper reached both exact URLs; local Windows curl then stopped at its CA-path configuration, so the before/after output comparison used the already-downloaded current HTML through the real DevDocs parser/filter stack.

Re-review is requested because the two issues identified from the screenshots are now covered directly.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants