Module:PageSummary:修订间差异
武外梗百科 爱国好学自强图新的百科全书
更多操作
无编辑摘要 |
无编辑摘要 |
||
| 第1行: | 第1行: | ||
local p = {} | local p = {} | ||
-- | -- 精准清理函数:保留加粗和链接,去除表格、模板和正文图片 | ||
local function | local function cleanContent(text) | ||
if not text then return "" end | if not text then return "" end | ||
-- 1. | -- 1. 移除 HTML 注释、脚注、数学公式 | ||
text = mw.ustring.gsub(text, "<!%-%-.-%-%->", "") | text = mw.ustring.gsub(text, "<!%-%-.-%-%->", "") | ||
text = mw.ustring.gsub(text, "<ref[^>]*>.-</ref>", "") | text = mw.ustring.gsub(text, "<ref[^>]*>.-</ref>", "") | ||
text = mw.ustring.gsub(text, "<ref[^>]-/>", "") | text = mw.ustring.gsub(text, "<ref[^>]-/>", "") | ||
-- 2. | -- 2. 移除表格 {| ... |} | ||
-- | -- 处理嵌套表格:从内向外剥离 | ||
text = mw.ustring.gsub(text, "{|[^{}]*|}", "") | local prev | ||
text = | repeat | ||
prev = text | |||
text = mw.ustring.gsub(text, "{|[^{}]*|}", "") | |||
until text == prev | |||
-- 3. | -- 3. 递归移除模板 {{ ... }} (保留正文文字,移除干扰块) | ||
local count = 0 | |||
local | |||
repeat | repeat | ||
prev = text | prev = text | ||
| 第25行: | 第26行: | ||
until text == prev or count > 10 | until text == prev or count > 10 | ||
-- 4. | -- 4. 【核心】去除正文中的图片/文件,但保留普通链接 | ||
-- 逻辑:匹配 [[...]],如果是文件则删掉,如果是链接则原样保留 | |||
text = mw.ustring.gsub(text, "%[%[([^%[%]]-)%]%]", function(inner) | text = mw.ustring.gsub(text, "%[%[([^%[%]]-)%]%]", function(inner) | ||
local | local low = mw.ustring.lower(inner) | ||
if | -- 识别文件、分类前缀 | ||
if low:match("^file:") or low:match("^image:") or | |||
low:match("^文件:") or low:match("^图像:") or | |||
return "" -- | low:match("^category:") or low:match("^分类:") then | ||
return "" -- 移除这些内容 | |||
end | end | ||
return "[[" .. inner .. "]]" -- 保留普通链接 | |||
end) | end) | ||
-- 5. | -- 5. 清理多余的标题符号、换行和空格 | ||
text = mw.ustring.gsub(text, "\n==+.-==+", " ") -- 移除标题行 | |||
text = mw.ustring.gsub(text, "==+.-==+", "") | text = mw.ustring.gsub(text, "\n%s*[*#:]+", " ") -- 列表转为空格 | ||
text = mw.ustring.gsub(text, "\n%s*[*#:]+", " ") -- | |||
text = mw.ustring.gsub(text, "\n+", " ") -- 换行转空格 | text = mw.ustring.gsub(text, "\n+", " ") -- 换行转空格 | ||
text = mw.ustring.gsub(text, "%s+", " ") -- | text = mw.ustring.gsub(text, "%s+", " ") -- 压缩空格 | ||
return mw.text.trim(text) | return mw.text.trim(text) | ||
| 第54行: | 第53行: | ||
if not title or not title.exists then return "" end | if not title or not title.exists then return "" end | ||
-- | -- 读取前 1500 字,保证性能的同时覆盖首段 | ||
local | local rawContent = title:getContent() or "" | ||
local limitedContent = mw.ustring.sub(rawContent, 1, 1500) | |||
-- 1. | -- 1. 提取缩略图(在清理前先抓取第一张图名) | ||
local firstImage = mw.ustring.match( | local firstImage = mw.ustring.match(limitedContent, "%[%[%s*[Ff]ile%s*:([^|%]%s]+)") or | ||
mw.ustring.match( | mw.ustring.match(limitedContent, "%[%[%s*文件%s*:([^|%]%s]+)") or | ||
mw.ustring.match( | mw.ustring.match(limitedContent, "%[%[%s*[Ii]mage%s*:([^|%]%s]+)") | ||
-- 2. | -- 2. 执行清理(保留加粗和链接) | ||
local cleanText = | local cleanText = cleanContent(limitedContent) | ||
local targetLen = | |||
-- 3. 截取摘要长度 | |||
local targetLen = 180 | |||
local summary = mw.ustring.sub(cleanText, 1, targetLen) | local summary = mw.ustring.sub(cleanText, 1, targetLen) | ||
-- 4. 补全因截断可能破坏的加粗或链接标签 | |||
if mw.ustring.len(cleanText) > targetLen then | if mw.ustring.len(cleanText) > targetLen then | ||
summary = summary .. "..." | summary = summary .. "..." | ||
-- 简单的标签闭合检查(防止截断导致的页面错乱) | |||
local _, opens = mw.ustring.gsub(summary, "'''", "") | |||
if opens % 2 ~= 0 then summary = summary .. "'''" end | |||
end | end | ||
-- | -- 5. 渲染(不加外框,仅保持文字环绕) | ||
local | local res = mw.html.create('div'):css({['display'] = 'flow-root'}) | ||
-- 渲染标题 | |||
res:tag('div') | |||
:css({['font-size'] = '1.2em', ['font-weight'] = 'bold', ['margin-bottom'] = '5px'}) | |||
:wikitext('[[' .. pageName .. ']]') | |||
-- | |||
-- | -- 渲染图片(缩略图) | ||
if firstImage then | if firstImage then | ||
res:wikitext('[[File:' .. firstImage .. '|120px|right|link=' .. pageName .. ']]') | |||
end | end | ||
-- 渲染带加粗和链接的正文 | |||
res:wikitext(summary) | |||
return tostring( | return tostring(res) | ||
end | end | ||
return p | return p | ||
2026年2月18日 (三) 13:29的版本
此模块的文档可以在Module:PageSummary/doc创建
local p = {}
-- 精准清理函数:保留加粗和链接,去除表格、模板和正文图片
local function cleanContent(text)
if not text then return "" end
-- 1. 移除 HTML 注释、脚注、数学公式
text = mw.ustring.gsub(text, "<!%-%-.-%-%->", "")
text = mw.ustring.gsub(text, "<ref[^>]*>.-</ref>", "")
text = mw.ustring.gsub(text, "<ref[^>]-/>", "")
-- 2. 移除表格 {| ... |}
-- 处理嵌套表格:从内向外剥离
local prev
repeat
prev = text
text = mw.ustring.gsub(text, "{|[^{}]*|}", "")
until text == prev
-- 3. 递归移除模板 {{ ... }} (保留正文文字,移除干扰块)
local count = 0
repeat
prev = text
text = mw.ustring.gsub(text, "{{[^{}]-}}", "")
count = count + 1
until text == prev or count > 10
-- 4. 【核心】去除正文中的图片/文件,但保留普通链接
-- 逻辑:匹配 [[...]],如果是文件则删掉,如果是链接则原样保留
text = mw.ustring.gsub(text, "%[%[([^%[%]]-)%]%]", function(inner)
local low = mw.ustring.lower(inner)
-- 识别文件、分类前缀
if low:match("^file:") or low:match("^image:") or
low:match("^文件:") or low:match("^图像:") or
low:match("^category:") or low:match("^分类:") then
return "" -- 移除这些内容
end
return "[[" .. inner .. "]]" -- 保留普通链接
end)
-- 5. 清理多余的标题符号、换行和空格
text = mw.ustring.gsub(text, "\n==+.-==+", " ") -- 移除标题行
text = mw.ustring.gsub(text, "\n%s*[*#:]+", " ") -- 列表转为空格
text = mw.ustring.gsub(text, "\n+", " ") -- 换行转空格
text = mw.ustring.gsub(text, "%s+", " ") -- 压缩空格
return mw.text.trim(text)
end
function p.getSummaryAndImage(frame)
local pageName = frame.args[1] or ""
local title = mw.title.new(pageName)
if not title or not title.exists then return "" end
-- 读取前 1500 字,保证性能的同时覆盖首段
local rawContent = title:getContent() or ""
local limitedContent = mw.ustring.sub(rawContent, 1, 1500)
-- 1. 提取缩略图(在清理前先抓取第一张图名)
local firstImage = mw.ustring.match(limitedContent, "%[%[%s*[Ff]ile%s*:([^|%]%s]+)") or
mw.ustring.match(limitedContent, "%[%[%s*文件%s*:([^|%]%s]+)") or
mw.ustring.match(limitedContent, "%[%[%s*[Ii]mage%s*:([^|%]%s]+)")
-- 2. 执行清理(保留加粗和链接)
local cleanText = cleanContent(limitedContent)
-- 3. 截取摘要长度
local targetLen = 180
local summary = mw.ustring.sub(cleanText, 1, targetLen)
-- 4. 补全因截断可能破坏的加粗或链接标签
if mw.ustring.len(cleanText) > targetLen then
summary = summary .. "..."
-- 简单的标签闭合检查(防止截断导致的页面错乱)
local _, opens = mw.ustring.gsub(summary, "'''", "")
if opens % 2 ~= 0 then summary = summary .. "'''" end
end
-- 5. 渲染(不加外框,仅保持文字环绕)
local res = mw.html.create('div'):css({['display'] = 'flow-root'})
-- 渲染标题
res:tag('div')
:css({['font-size'] = '1.2em', ['font-weight'] = 'bold', ['margin-bottom'] = '5px'})
:wikitext('[[' .. pageName .. ']]')
-- 渲染图片(缩略图)
if firstImage then
res:wikitext('[[File:' .. firstImage .. '|120px|right|link=' .. pageName .. ']]')
end
-- 渲染带加粗和链接的正文
res:wikitext(summary)
return tostring(res)
end
return p