プロジェクト

全般

プロフィール

Vote #65602

未完了

add some additional URL paths to robots.txt

Admin Redmine さんが約4年前に追加. 約4年前に更新.

ステータス:
New
優先度:
通常
担当者:
-
カテゴリ:
-
対象バージョン:
-
開始日:
2009/08/18
期日:
進捗率:

0%

予定工数:
category_id:
0
version_id:
0
issue_org_id:
3754
author_id:
7480
assigned_to_id:
0
comments:
10
status_id:
1
tracker_id:
3
plus1:
0
affected_version:
closed_on:
affected_version_id:
ステータス-->[New]

説明

My apache logs show that some redmine URLs are being heavily indexed by robots, and it seems like it would be best to have them blocked by robots.txt:
/issues
/projects//time_entries
/projects/N/wiki/
(where N is the numeric project id)
/repositories/annotate/*
/repositories/browse/*
/repositories/changes/*
/repositories/diff/*
/repositories/entry/*


journals

See the [[Plugin_List#Bots-Filter-plugin|Bots Filter plugin]] which has some overlap (e.g. the repositories). Maybe you can modify it to adapt it to your precise requirements?

Regards,

Mischa.
--------------------------------------------------------------------------------
Or I can easily block these via my apache config. But I do think they should be added to robots.txt by default. I also wonder, how are Googlebot and others even finding some of these non-canonical paths? It could point to a bug elsewhere which is generating links to these paths?
--------------------------------------------------------------------------------
Here's a patch adding the additional problematic paths to the default robots.txt
--------------------------------------------------------------------------------
I like having the robots crawl some of these pages, they even turn up when I'm searching for a bug that I've already fixed.

* wiki pages
* global issues list
* repositories
--------------------------------------------------------------------------------
The wiki pages that this patch blocks are not the canonical path, they use the numeric project id rather than project name.

I now realize that the initial version of this patch blocked the individual issue pages; I intended to only block /issues? -- i.e. the global issue search page.
--------------------------------------------------------------------------------

--------------------------------------------------------------------------------
My site is also getting hammer on /repositories and /issues. Seems somewhat pointless to disallow access to these resources through /projects/... but not other urls.
--------------------------------------------------------------------------------
This patch has been ready for more than 3 years, why hasn't this been committed yet?
--------------------------------------------------------------------------------
Here's an updated patch for 1.4.
--------------------------------------------------------------------------------
Antoine Beaupré wrote:
> Here's an updated patch for 1.4.

and that was now two years ago, with the patch sitting here for 5 years. can we at least get feedback on what's wrong with the patch, if anything?

thanks.
--------------------------------------------------------------------------------


related_issues

relates,Closed,6734,robots.txt: disallow crawling issues list with a query string

表示するデータがありません

他の形式にエクスポート: Atom PDF

いいね!0