# Limitations

The laboratory is intentionally small and local. Agent suites are modest, code tasks are single-bug and primarily single-file, keyword correctness is coarse, and local model reruns can vary. Foundation measurements depend on the original external training setup. No result here should be generalized beyond its stated hardware, software, configuration, and task set.

Deep research discovers general-web results through a bounded no-key Bing RSS
surface and scholarly metadata/abstracts through Semantic Scholar, arXiv, and
Crossref. Web extraction handles static public HTTPS HTML and plain text; it
does not execute JavaScript, authenticate, bypass paywalls, defeat bot
challenges, or ignore robots.txt. Search quality and availability still depend
on upstream providers. Citation-ID and content-hash checks prevent the model
from inventing references outside the collected set or citing unextracted and
prompt-injection-shaped webpages, but semantic claim verification and
publication decisions still require human review of the primary source.
