# As a condition of accessing this website, you agree to abide by the following # content signals: # (a) If a Content-Signal = yes, you may collect content for the corresponding # use. # (b) If a Content-Signal = no, you may not collect content for the # corresponding use. # (c) If the website operator does not include a Content-Signal for a # corresponding use, the website operator neither grants nor restricts # permission via Content-Signal with respect to the corresponding use. # The content signals and their meanings are: # search: building a search index and providing search results (e.g., returning # hyperlinks and short excerpts from your website's contents). Search does not # include providing AI-generated search summaries. # ai-input: inputting content into one or more AI models (e.g., retrieval # augmented generation, grounding, or other real-time taking of content for # generative AI search answers). # ai-train: training or fine-tuning AI models. # use: how AI systems may consume the content (immediate, reference, or full). # ANY RESTRICTIONS EXPRESSED VIA CONTENT SIGNALS ARE EXPRESS RESERVATIONS OF # RIGHTS UNDER ARTICLE 4 OF THE EUROPEAN UNION DIRECTIVE 2019/790 ON COPYRIGHT # AND RELATED RIGHTS IN THE DIGITAL SINGLE MARKET. # BEGIN Cloudflare Managed content User-agent: * Content-Signal: search=yes,ai-train=no,use=reference Allow: / User-agent: Amazonbot Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: Bytespider Disallow: / User-agent: CCBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: CloudflareBrowserRenderingCrawler Disallow: / User-agent: Google-Extended Disallow: / User-agent: GPTBot Disallow: / User-agent: meta-externalagent Disallow: / # END Cloudflare Managed Content # 진짜 일본어 — robots.txt # # ── Content Signals ─────────────────────────────────────── # Disallow 가 「긁지 마라」라면, 이건 「긁은 다음 어디에 쓰지 마라」다. 축이 셋이다. # search = 검색 색인·링크·짧은 인용 # ai-input = 실시간 생성 답변에 인용(RAG) # ai-train = 모델 학습·파인튜닝 # # 인용은 방문자를 데려오지만 학습은 데이터만 가져간다. 그래서 앞의 둘은 열고 학습만 막는다. # # 이 문서의 신호를 따르는 것은 그 신호를 존중하겠다는 표시이며, 무시하고 수집·학습에 # 사용하는 것은 아래 이용약관 위반이다. — /terms.html Content-Signal: search=yes, ai-input=yes, ai-train=no # # 판단 기준은 crawl-to-refer 비율이다. 몇 페이지를 긁고 방문자를 몇 명 보내주는가. # · 학습용 크롤러 — 유입 0, 데이터만 가져간다 → 막는다 # · 검색·인용 봇 — 인용을 통해 방문자를 보낸다 → 연다 # # ⚠️ 회사마다 봇 이름이 여러 개로 갈려 있다. 하나만 막으면 나머지는 그대로 들어온다. # (ClaudeBot 만 막으면 Claude-SearchBot·Claude-User 는 통과) # ── 학습용 — 차단 ────────────────────────────────────────── User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: anthropic-ai Disallow: / User-agent: CCBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: FacebookBot Disallow: / User-agent: Bytespider Disallow: / User-agent: Amazonbot Disallow: / User-agent: cohere-ai Disallow: / User-agent: Diffbot Disallow: / User-agent: Omgilibot Disallow: / User-agent: PetalBot Disallow: / User-agent: Timpibot Disallow: / User-agent: ImagesiftBot Disallow: / # ── 검색·인용 — 허용하되 API 는 막는다 ─────────────────── # 인용을 통한 유입은 받는다. 다만 JSON 엔드포인트를 긁을 이유는 없다. User-agent: OAI-SearchBot Disallow: /api/ User-agent: ChatGPT-User Disallow: /api/ User-agent: Claude-SearchBot Disallow: /api/ User-agent: Claude-User Disallow: /api/ User-agent: PerplexityBot Disallow: /api/ User-agent: Perplexity-User Disallow: /api/ # ── 일반 검색엔진 ──────────────────────────────────────── User-agent: * Disallow: /api/ Crawl-delay: 2 # 예문 데이터는 수작업 검수의 결과물이다. 자동 수집·재배포를 금지한다. # 이용약관: /terms.html Sitemap: https://realjapanese.kr/sitemap.xml