Module:PageSummary:修订间差异
武外梗百科 爱国好学自强图新的百科全书
更多操作
无编辑摘要 |
无编辑摘要 |
||
| 第1行: | 第1行: | ||
local p = {} | local p = {} | ||
-- | -- 强化版清理:彻底根除内文图片,保留加粗和链接 | ||
local function cleanContent(text) | local function cleanContent(text) | ||
if not text then return "" end | if not text then return "" end | ||
-- 1. | -- 1. 基础清理(注释、脚注、表格、模板) | ||
text = mw.ustring.gsub(text, "<!%-%-.-%-%->", "") | text = mw.ustring.gsub(text, "<!%-%-.-%-%->", "") | ||
text = mw.ustring.gsub(text, "<ref[^>]*>.-</ref>", "") | text = mw.ustring.gsub(text, "<ref[^>]*>.-</ref>", "") | ||
text = mw.ustring.gsub(text, "<ref[^>]-/>", "") | text = mw.ustring.gsub(text, "<ref[^>]-/>", "") | ||
-- | -- 移除表格 | ||
local prev | local prev | ||
repeat | repeat | ||
| 第18行: | 第17行: | ||
until text == prev | until text == prev | ||
-- | -- 移除模板 | ||
local count = 0 | local count = 0 | ||
repeat | repeat | ||
| 第26行: | 第25行: | ||
until text == prev or count > 10 | until text == prev or count > 10 | ||
-- | -- 2. 【强化】彻底剔除正文中的图片引用 | ||
-- | -- 这里的正则覆盖了包含换行、管道符、嵌套括号的复杂情况 | ||
text = mw.ustring.gsub(text, "%[%[([^%[%]]-)%]%]", function(inner) | -- 我们使用递归匹配,确保嵌套的 [[文件:[[链接]]]] 也能被干掉 | ||
repeat | |||
prev = text | |||
text = mw.ustring.gsub(text, "%[%[%s*([^%[%]]-)%s*%]%]", function(inner) | |||
local low = mw.ustring.lower(inner) | |||
-- 如果是文件、分类、图像等,返回空字符串(彻底剔除) | |||
if low:match("^file:") or low:match("^image:") or | |||
low:match("^文件:") or low:match("^图像:") or | |||
low:match("^category:") or low:match("^分类:") then | |||
return "" | |||
end | |||
-- 如果是普通链接,暂时保留 | |||
return "%%LINKSTART%%" .. inner .. "%%LINKEND%%" | |||
end) | |||
until text == prev | |||
-- 3. 清理剩余的 Wikitext 杂质 | |||
text = mw.ustring.gsub(text, "\n==+.-==+", " ") | |||
text = mw.ustring.gsub(text, "\n%s*[*#:]+", " ") | |||
text = mw.ustring.gsub(text, "\n+", " ") | |||
text = mw.ustring.gsub(text, "%s+", " ") | |||
-- | -- 4. 恢复链接(将占位符转回标准链接) | ||
text = mw.ustring.gsub(text, " | text = mw.ustring.gsub(text, "%%LINKSTART%%(.-)%%LINKEND%%", "[[%1]]") | ||
return mw.text.trim(text) | return mw.text.trim(text) | ||
| 第53行: | 第60行: | ||
if not title or not title.exists then return "" end | if not title or not title.exists then return "" end | ||
local rawContent = title:getContent() or "" | local rawContent = title:getContent() or "" | ||
local limitedContent = mw.ustring.sub(rawContent, 1, | -- 只读取前 2000 字符,确保护盖首段图片 | ||
local limitedContent = mw.ustring.sub(rawContent, 1, 2000) | |||
-- 1. | -- 1. 提取缩略图(在彻底清理前提取出文件名) | ||
local firstImage = mw.ustring.match(limitedContent, "%[%[%s*[Ff]ile%s*:([^|%]%s]+)") or | local firstImage = mw.ustring.match(limitedContent, "%[%[%s*[Ff]ile%s*:([^|%]%s\n]+)") or | ||
mw.ustring.match(limitedContent, "%[%[%s*文件%s*:([^|%]%s]+)") or | mw.ustring.match(limitedContent, "%[%[%s*文件%s*:([^|%]%s\n]+)") or | ||
mw.ustring.match(limitedContent, "%[%[%s*[Ii]mage%s*:([^|%]%s]+)") | mw.ustring.match(limitedContent, "%[%[%s*[Ii]mage%s*:([^|%]%s\n]+)") | ||
-- 2. | -- 2. 执行强力清理 | ||
local cleanText = cleanContent(limitedContent) | local cleanText = cleanContent(limitedContent) | ||
-- 3. | -- 3. 安全截断(处理中文和标签闭合) | ||
local targetLen = | local targetLen = 150 | ||
local summary = | local summary = "" | ||
if mw.ustring.len(cleanText) <= targetLen then | |||
if mw.ustring.len(cleanText) | summary = cleanText | ||
summary = summary .. "..." | else | ||
-- | -- 使用 ustring 截断防止乱码,并加上省略号 | ||
summary = mw.ustring.sub(cleanText, 1, targetLen) .. "..." | |||
-- 闭合加粗标签 | |||
local _, opens = mw.ustring.gsub(summary, "'''", "") | local _, opens = mw.ustring.gsub(summary, "'''", "") | ||
if opens % 2 ~= 0 then summary = summary .. "'''" end | if opens % 2 ~= 0 then summary = summary .. "'''" end | ||
-- 闭合链接标签(如果截断在链接中间,直接舍弃最后一部分) | |||
if mw.ustring.match(summary, "%[%[[^%]]*$") then | |||
summary = mw.ustring.gsub(summary, "%[%[[^%]]*$", "") .. "..." | |||
end | |||
end | end | ||
-- | -- 4. 最终渲染 | ||
local res = mw.html.create('div'):css({['display'] = 'flow-root'}) | local res = mw.html.create('div'):css({['display'] = 'flow-root', ['line-height'] = '1.6'}) | ||
-- | -- 标题 | ||
res:tag('div') | res:tag('div') | ||
:css({['font-size'] = '1.2em', ['font-weight'] = 'bold', ['margin-bottom'] = '5px'}) | :css({['font-size'] = '1.2em', ['font-weight'] = 'bold', ['margin-bottom'] = '5px'}) | ||
:wikitext('[[' .. pageName .. ']]') | :wikitext('[[' .. pageName .. ']]') | ||
-- | -- 图片 | ||
if firstImage then | if firstImage then | ||
res:wikitext('[[File:' .. firstImage .. '|120px|right|link=' .. pageName .. ']]') | res:wikitext('[[File:' .. firstImage .. '|120px|right|link=' .. pageName .. ']]') | ||
end | end | ||
-- | -- 正文(带加粗和链接) | ||
res:wikitext(summary) | res:wikitext(summary) | ||
2026年2月18日 (三) 13:31的版本
此模块的文档可以在Module:PageSummary/doc创建
local p = {}
-- 强化版清理:彻底根除内文图片,保留加粗和链接
local function cleanContent(text)
if not text then return "" end
-- 1. 基础清理(注释、脚注、表格、模板)
text = mw.ustring.gsub(text, "<!%-%-.-%-%->", "")
text = mw.ustring.gsub(text, "<ref[^>]*>.-</ref>", "")
text = mw.ustring.gsub(text, "<ref[^>]-/>", "")
-- 移除表格
local prev
repeat
prev = text
text = mw.ustring.gsub(text, "{|[^{}]*|}", "")
until text == prev
-- 移除模板
local count = 0
repeat
prev = text
text = mw.ustring.gsub(text, "{{[^{}]-}}", "")
count = count + 1
until text == prev or count > 10
-- 2. 【强化】彻底剔除正文中的图片引用
-- 这里的正则覆盖了包含换行、管道符、嵌套括号的复杂情况
-- 我们使用递归匹配,确保嵌套的 [[文件:[[链接]]]] 也能被干掉
repeat
prev = text
text = mw.ustring.gsub(text, "%[%[%s*([^%[%]]-)%s*%]%]", function(inner)
local low = mw.ustring.lower(inner)
-- 如果是文件、分类、图像等,返回空字符串(彻底剔除)
if low:match("^file:") or low:match("^image:") or
low:match("^文件:") or low:match("^图像:") or
low:match("^category:") or low:match("^分类:") then
return ""
end
-- 如果是普通链接,暂时保留
return "%%LINKSTART%%" .. inner .. "%%LINKEND%%"
end)
until text == prev
-- 3. 清理剩余的 Wikitext 杂质
text = mw.ustring.gsub(text, "\n==+.-==+", " ")
text = mw.ustring.gsub(text, "\n%s*[*#:]+", " ")
text = mw.ustring.gsub(text, "\n+", " ")
text = mw.ustring.gsub(text, "%s+", " ")
-- 4. 恢复链接(将占位符转回标准链接)
text = mw.ustring.gsub(text, "%%LINKSTART%%(.-)%%LINKEND%%", "[[%1]]")
return mw.text.trim(text)
end
function p.getSummaryAndImage(frame)
local pageName = frame.args[1] or ""
local title = mw.title.new(pageName)
if not title or not title.exists then return "" end
local rawContent = title:getContent() or ""
-- 只读取前 2000 字符,确保护盖首段图片
local limitedContent = mw.ustring.sub(rawContent, 1, 2000)
-- 1. 提取缩略图(在彻底清理前提取出文件名)
local firstImage = mw.ustring.match(limitedContent, "%[%[%s*[Ff]ile%s*:([^|%]%s\n]+)") or
mw.ustring.match(limitedContent, "%[%[%s*文件%s*:([^|%]%s\n]+)") or
mw.ustring.match(limitedContent, "%[%[%s*[Ii]mage%s*:([^|%]%s\n]+)")
-- 2. 执行强力清理
local cleanText = cleanContent(limitedContent)
-- 3. 安全截断(处理中文和标签闭合)
local targetLen = 150
local summary = ""
if mw.ustring.len(cleanText) <= targetLen then
summary = cleanText
else
-- 使用 ustring 截断防止乱码,并加上省略号
summary = mw.ustring.sub(cleanText, 1, targetLen) .. "..."
-- 闭合加粗标签
local _, opens = mw.ustring.gsub(summary, "'''", "")
if opens % 2 ~= 0 then summary = summary .. "'''" end
-- 闭合链接标签(如果截断在链接中间,直接舍弃最后一部分)
if mw.ustring.match(summary, "%[%[[^%]]*$") then
summary = mw.ustring.gsub(summary, "%[%[[^%]]*$", "") .. "..."
end
end
-- 4. 最终渲染
local res = mw.html.create('div'):css({['display'] = 'flow-root', ['line-height'] = '1.6'})
-- 标题
res:tag('div')
:css({['font-size'] = '1.2em', ['font-weight'] = 'bold', ['margin-bottom'] = '5px'})
:wikitext('[[' .. pageName .. ']]')
-- 图片
if firstImage then
res:wikitext('[[File:' .. firstImage .. '|120px|right|link=' .. pageName .. ']]')
end
-- 正文(带加粗和链接)
res:wikitext(summary)
return tostring(res)
end
return p