正则表达式的奇技淫巧

需求

在做爬虫的时候,需要对下面的链接进行字段替换

python中在不同的阶段要做不同的处理,不管是用 f""  或是format都不太好处理, 在不同的阶段要做不同的处理 

URL = 'https://www.abc.com/order/index?filter={"type":0,"user":"","create_time":"2024-01-01 00:00:00 - 2024-01-01 23:59:59","item_title":"","status":"0","adzone_id":"","time":"2","media_id":""}&sort=id&order=asc&offset=0&limit=100'

解决方案

自己写一个模板替换函数yyds

python版本

def format(str, pattern=r"##(.*?)##", **kwargs):
    """
    模板替换字符串
    """
    res = re.sub(pattern, lambda x: kwargs.get(x.group(1),x.group()), str)
    return res
    
    
URL = 'https://www.abc.com/order/index?filter={"type":0,"user":"","create_time":"##create_time##","item_title":"","status":"0","adzone_id":"","time":"2","media_id":""}&sort=id&order=asc&offset=##offset##&limit=##limit##'
# 第一次进行替换
url = format(URL, create_time="2024-01-01 00:00:00 - 2024-01-01 23:59:59")
print(url)
# 'https://www.abc.com/order/index?filter={"type":0,"user":"","create_time":"2024-01-01 00:00:00 - 2024-01-01 23:59:59","item_title":"","status":"0","adzone_id":"","time":"2","media_id":""}&sort=id&order=asc&offset=##offset##&limit=##limit##'
# 第二次进行替换
paginates = {"offset": "0", "limit": "100"}
url = format(url, **paginates)
print(url)
# 'https://www.abc.com/order/index?filter={"type":0,"user":"","create_time":"2024-01-01 00:00:00 - 2024-01-01 23:59:59","item_title":"","status":"0","adzone_id":"","time":"2","media_id":""}&sort=id&order=asc&offset=0&limit=100'

javascript版本

function format(str, key_map, pattern){
	pattern = pattern || /##(.*?)##/g
	res = str.replace(pattern, (match, group)=>{return key_map[group]||match});
	return res
}
const URL = 'https://www.abc.com/order/index?filter={"type":0,"user":"","create_time":"##create_time##","item_title":"","status":"0","adzone_id":"","time":"2","media_id":""}&sort=id&order=asc&offset=##offset##&limit=##limit##'
// 第一次进行替换
let url = format(URL, {create_time: "2024-01-01 00:00:00 - 2024-01-01 23:59:59"})
console.log(url)
// 'https://www.abc.com/order/index?filter={"type":0,"user":"","create_time":"2024-01-01 00:00:00 - 2024-01-01 23:59:59","item_title":"","status":"0","adzone_id":"","time":"2","media_id":""}&sort=id&order=asc&offset=##offset##&limit=##limit##'
# 第二次进行替换
let paginates = {"offset": "0", "limit": "100"}
let url = format(url, paginates)
console.log(url)
# 'https://www.abc.com/order/index?filter={"type":0,"user":"","create_time":"2024-01-01 00:00:00 - 2024-01-01 23:59:59","item_title":"","status":"0","adzone_id":"","time":"2","media_id":""}&sort=id&order=asc&offset=0&limit=100'

查看原文:正则表达式的奇技淫巧

关注公众号 "字节航海家" 及时获取最新内容

  • 18
    点赞
  • 9
    收藏
    觉得还不错? 一键收藏
  • 0
    评论
评论
添加红包

请填写红包祝福语或标题

红包个数最小为10个

红包金额最低5元

当前余额3.43前往充值 >
需支付:10.00
成就一亿技术人!
领取后你会自动成为博主和红包主的粉丝 规则
hope_wisdom
发出的红包
实付
使用余额支付
点击重新获取
扫码支付
钱包余额 0

抵扣说明:

1.余额是钱包充值的虚拟货币,按照1:1的比例进行支付金额的抵扣。
2.余额无法直接购买下载,可以购买VIP、付费专栏及课程。

余额充值