11赞

从Python中删除字符串标点符号的最佳方法

作者：周扒pi | 2023-09-03 11:31

如何解决《从Python中删除字符串标点符号的最佳方法》经验，为你挑选了15个好方法。

似乎应该有一个比以下更简单的方法:

import string
s = "string. With. Punctuation?" # Sample string 
out = s.translate(string.maketrans("",""), string.punctuation)

在那儿？

1> Brian..：

从效率的角度来看,你不会打败

s.translate(None, string.punctuation)

它使用查找表在C中执行原始字符串操作 - 除了编写自己的C代码之外,没有什么能比这更好.

如果速度不是担心,另一个选择是:

s.translate(str.maketrans('', '', string.punctuation))

这比使用每个char的s.replace更快,但是不能像非纯python方法那样执行,例如regexes或string.translate,正如您可以从下面的时间看到的那样.对于这种类型的问题,尽可能低的水平做到这一点是值得的.

时间码:

exclude = set(string.punctuation)
s = ''.join(ch for ch in s if ch not in exclude)

这给出了以下结果:

import re, string, timeit

s = "string. With. Punctuation"
exclude = set(string.punctuation)
table = string.maketrans("","")
regex = re.compile('[%s]' % re.escape(string.punctuation))

def test_set(s):
    return ''.join(ch for ch in s if ch not in exclude)

def test_re(s):  # From Vinko's solution, with fix.
    return regex.sub('', s)

def test_trans(s):
    return s.translate(table, string.punctuation)

def test_repl(s):  # From S.Lott's solution
    for c in string.punctuation:
        s=s.replace(c,"")
    return s

print "sets      :",timeit.Timer('f(s)', 'from __main__ import s,test_set as f').timeit(1000000)
print "regex     :",timeit.Timer('f(s)', 'from __main__ import s,test_re as f').timeit(1000000)
print "translate :",timeit.Timer('f(s)', 'from __main__ import s,test_trans as f').timeit(1000000)
print "replace   :",timeit.Timer('f(s)', 'from __main__ import s,test_repl as f').timeit(1000000)

在Python3中,`table = string.maketrans("","")`应该替换为`table = str.maketrans({key:None for string in string.punctuation})`？

很好的答案.您可以通过删除表来简化它.文档说:"对于只删除字符的翻译,将表格参数设置为无"(http://docs.python.org/library/stdtypes.html#str.translate)

感谢时间信息,我正在考虑自己做类似的事情,但你的写作比我想做的更好,现在我可以用它作为我想写的任何未来时序代码的模板:).

要更新讨论,从Python 3.6开始,`regex`现在是最有效的方法!它几乎比翻译快2倍.而且,设置和替换不再那么糟糕!他们都提高了4倍:)

值得注意的是,translate()对str和unicode对象的行为有所不同,所以你需要确保你总是使用相同的数据类型,但这个答案中的方法同样适用于两者,这很方便.

@mlissner - 效率.它是一个列表/字符串,您需要进行线性扫描以确定字母是否在字符串中.但是,使用集合或字典,它通常会更快(除了非常小的字符串),因为它不必检查每个值.

在Python 3中，转换表也可以通过`table = str.maketrans（''，``，string.punctuation）`https://docs.python.org/3/library/stdtypes.html#str创建。 maketrans

2> Eratosthenes..：

正则表达式很简单,如果你知道的话.

import re
s = "string. With. Punctuation?"
s = re.sub(r'[^\w\s]','',s)

在上面的代码中,我们用空字符串替换(re.sub)所有NON [字母数字字符(\ w)和空格(\ s)].
因此.和？通过正则表达式运行s变量后,变量's'中不会出现标点符号.

@Outlier说明:用空字符串替换不是(^)字符或空格.但要小心,\ w通常也会匹配下划线.

@SIslam我认为它将与unicode一起使用unicode标志设置,即`s = re.sub(r'[^\w\s]','',s,re.UNICODE)`.在linux上使用python 3进行测试即使没有使用tamil字母,தமிழ்的标志也能正常工作.

3> SparkAndShin..：

为了方便使用,我总结了Python 2和Python 3中字符串条带标点符号的注释.请参阅其他答案以获取详细说明.

Python 2

import string

s = "string. With. Punctuation?"
table = string.maketrans("","")
new_s = s.translate(table, string.punctuation)      # Output: string without punctuation

Python 3

import string

s = "string. With. Punctuation?"
table = str.maketrans(dict.fromkeys(string.punctuation))  # OR {key: None for key in string.punctuation}
new_s = s.translate(table)                          # Output: string without punctuation

4> 小智..：

myString.translate(None, string.punctuation)

`TypeError:translate()只取一个参数(给定2个)`:(

请注意,对于Python 3中的`str`和Python 2中的`unicode`,不支持`deletechars`参数.

啊,我试过这个,但它并不适用于所有情况.myString.translate(string.maketrans("",""),string.punctuation)工作正常.

myString.translate(string.maketrans("",""),string.punctuation)不适用于unicode字符串(找出困难的方法)

@BrianTingle:在我的评论中查看Python 3代码(它传递一个参数).[按照链接,查看适用于unicode的Python 2代码](http://stackoverflow.com/a/11066687/4279)和[它的Python 3改编](http://stackoverflow.com/a/21635971/ 4279)

@agf:你仍然可以[使用`.translate()`删除标点符号,即使在Unicode和py3k情况下也是如此](http://stackoverflow.com/a/11066687/4279)使用字典参数.

@MarcMaxson：myString.translate（str.maketrans（“”，“”，string.punctuation））确实适用于Python 3上的Unicode字符串。尽管string.punctuation`仅包含ASCII标点。单击[我以前的评论中的链接]（http://stackoverflow.com/a/11066687/4279）。它显示了如何删除所有标点符号（包括Unicode标点符号）。

5> S.Lott..：

我经常使用这样的东西:

>>> s = "string. With. Punctuation?" # Sample string
>>> import string
>>> for c in string.punctuation:
...     s= s.replace(c,"")
...
>>> s
'string With Punctuation'

uglified one-liner:`reduce(lambda s,c:s.replace(c,''),string.punctuation,s)`.

6> Björn Lindqv..：

string.punctuation是ASCII 只!更正确(但也更慢)的方法是使用unicodedata模块:

# -*- coding: utf-8 -*-
from unicodedata import category
s = u'String — with -  «punctation »...'
s = ''.join(ch for ch in s if category(ch)[0] != 'P')
print 'stripped', s

你可以:[`regex.sub(ur"\ p {P} +","",text)`](http://stackoverflow.com/a/11066687/4279)

7> Vinko Vrsalo..：

如果你对家庭更熟悉,不一定更简单,但不一样.

import re, string
s = "string. With. Punctuation?" # Sample string 
out = re.sub('[%s]' % re.escape(string.punctuation), '', s)

实际上,它还是错的.序列"\\]"被视为一个转义(巧合地没有关闭],因此绕过了另一个失败),但是离开了\ unescaped.您应该使用re.escape(string.punctuation)来防止这种情况.

8> Martijn Piet..：

对于Python 3 str或Python 2 unicode值,str.translate()只需要一个字典; 在该映射中查找代码点(整数),并None删除映射到的任何内容.

要删除(某些？)标点符号,请使用:

import string

remove_punct_map = dict.fromkeys(map(ord, string.punctuation))
s.translate(remove_punct_map)

所述dict.fromkeys()类方法使得它琐碎创建映射,所有的值设置为None基于密钥的序列.

要删除所有标点符号,而不仅仅是ASCII标点符号,您的表格需要更大一些; 请参阅JF Sebastian的回答(Python 3版本):

import unicodedata
import sys

remove_punct_map = dict.fromkeys(i for i in range(sys.maxunicode)
                                 if unicodedata.category(chr(i)).startswith('P'))

9> Zach..：

string.punctuation错过了现实世界中常用的标点符号.如何使用适用于非ASCII标点符号的解决方案？

import regex
s = u"string. With. Some?Really Weird?Non?ASCII? ??Punctuation???"
remove = regex.compile(ur'[\p{C}|\p{M}|\p{P}|\p{S}|\p{Z}]+', regex.UNICODE)
remove.sub(u" ", s).strip()

就个人而言,我认为这是从Python中删除字符串标点符号的最佳方法,因为:

它删除所有Unicode标点符号

它很容易修改,例如你可以删除\{S}如果你想删除标点,但保持符号$.

您可以非常具体地了解要保留的内容以及要删除的内容,例如,\{Pd}只删除短划线.

这个正则表达式也规范了空白.它将标签,回车和其他奇怪的地方映射到漂亮的单个空间.

这使用Unicode字符属性,您可以在维基百科上阅读更多信息.

10> 小智..：

这是Python 3.5的单行程序:

import string
"l*ots! o(f. p@u)n[c}t]u[a'ti\"on#$^?/".translate(str.maketrans({a:None for a in string.punctuation}))

11> Blairg23..：

我还没有看到这个答案.只需使用正则表达式; 它除了单词字符(\w)和数字字符(\d)之外的所有字符,后跟一个空白字符(\s):

import re
s = "string. With. Punctuation?" # Sample string 
out = re.sub(ur'[^\w\d\s]+', '', s)

12> 小智..：

这可能不是最好的解决方案,但这就是我做到的.

import string
f = lambda x: ''.join([i for i in x if i not in string.punctuation])

13> Dr.Tautology..：

这是我写的一个函数.它不是很有效,但它很简单,您可以添加或删除任何您想要的标点符号:

def stripPunc(wordList):
    """Strips punctuation from list of words"""
    puncList = [".",";",":","!","?","/","\\",",","#","@","$","&",")","(","\""]
    for punc in puncList:
        for word in wordList:
            wordList=[word.replace(punc,'') for word in wordList]
    return wordList

14> Haythem HADH..：

import re
s = "string. With. Punctuation?" # Sample string 
out = re.sub(r'[^a-zA-Z0-9\s]', '', s)

15> krinker..：

作为更新，我重写了Python 3中的@Brian示例，并对其进行了更改，以将regex编译步骤移至函数内部。我的想法是计时使该功能起作用所需的每个步骤。也许您使用的是分布式计算，并且您的工作人员之间无法共享正则表达式对象，因此需要re.compile在每个工作人员中走一步。另外，我很好奇地为Python 3的maketrans的两种不同实现计时了

table = str.maketrans({key: None for key in string.punctuation})

与

table = str.maketrans('', '', string.punctuation)

另外，我添加了另一种使用set的方法，其中利用了交集函数来减少迭代次数。

这是完整的代码：

import re, string, timeit

s = "string. With. Punctuation"


def test_set(s):
    exclude = set(string.punctuation)
    return ''.join(ch for ch in s if ch not in exclude)


def test_set2(s):
    _punctuation = set(string.punctuation)
    for punct in set(s).intersection(_punctuation):
        s = s.replace(punct, ' ')
    return ' '.join(s.split())


def test_re(s):  # From Vinko's solution, with fix.
    regex = re.compile('[%s]' % re.escape(string.punctuation))
    return regex.sub('', s)


def test_trans(s):
    table = str.maketrans({key: None for key in string.punctuation})
    return s.translate(table)


def test_trans2(s):
    table = str.maketrans('', '', string.punctuation)
    return(s.translate(table))


def test_repl(s):  # From S.Lott's solution
    for c in string.punctuation:
        s=s.replace(c,"")
    return s


print("sets      :",timeit.Timer('f(s)', 'from __main__ import s,test_set as f').timeit(1000000))
print("sets2      :",timeit.Timer('f(s)', 'from __main__ import s,test_set2 as f').timeit(1000000))
print("regex     :",timeit.Timer('f(s)', 'from __main__ import s,test_re as f').timeit(1000000))
print("translate :",timeit.Timer('f(s)', 'from __main__ import s,test_trans as f').timeit(1000000))
print("translate2 :",timeit.Timer('f(s)', 'from __main__ import s,test_trans2 as f').timeit(1000000))
print("replace   :",timeit.Timer('f(s)', 'from __main__ import s,test_repl as f').timeit(1000000))

这是我的结果：

sets      : 3.1830138750374317
sets2      : 2.189873124472797
regex     : 7.142953420989215
translate : 4.243278483860195
translate2 : 2.427158243022859
replace   : 4.579746678471565

推荐阅读

程序员
Scikit-learn zip参数#1必须支持迭代

如何解决《Scikit-learnzip参数#1必须支持迭代》经验，为你挑选了1个好方法。 ... [详细]
程序员
在excel中剪切字符串

如何解决《在excel中剪切字符串》经验，为你挑选了1个好方法。 ... [详细]
程序员
禁用Jupyter键盘快捷键

如何解决《禁用Jupyter键盘快捷键》经验，为你挑选了0个好方法。 ... [详细]
程序员
使用变量中的字典作为函数的一组参数

如何解决《使用变量中的字典作为函数的一组参数》经验，为你挑选了1个好方法。 ... [详细]
程序员
在AWS Elastic Beanstalk上使用Django - Oscar自动设置Apache Solr

如何解决《在AWSElasticBeanstalk上使用Django-Oscar自动设置ApacheSolr》经验，为你挑选了0个好方法。 ... [详细]
程序员
获取一个数组并为每个项添加一个值？

如何解决《获取一个数组并为每个项添加一个值？》经验，为你挑选了1个好方法。 ... [详细]
程序员
Parsed_concat_1 @ 02ad98e0无法在Parsed_concat_1 ffmpeg上配置输出面板

如何解决《Parsed_concat_1@02ad98e0无法在Parsed_concat_1ffmpeg上配置输出面板》经验，为你挑选了1个好方法。 ... [详细]
程序员
在场景大纲之前运行一次特定步骤 - Python Behave

如何解决《在场景大纲之前运行一次特定步骤-PythonBehave》经验，为你挑选了0个好方法。 ... [详细]
程序员
将第一个列表的元素分配给另一个列表/数组

如何解决《将第一个列表的元素分配给另一个列表/数组》经验，为你挑选了1个好方法。 ... [详细]
程序员
如何引用其他.jsx文件中定义的组件

如何解决《如何引用其他.jsx文件中定义的组件》经验，为你挑选了1个好方法。 ... [详细]
程序员
如何在bash脚本中使用'history-c'命令？

如何解决《如何在bash脚本中使用'history-c'命令？》经验，为你挑选了0个好方法。 ... [详细]
程序员
在laravel刀片模板中爆炸字符串

如何解决《在laravel刀片模板中爆炸字符串》经验，为你挑选了1个好方法。 ... [详细]
程序员
rbenv:未安装版本"2.2.3"(由RBENV_VERSION环境变量设置)

如何解决《rbenv:未安装版本"2.2.3"(由RBENV_VERSION环境变量设置)》经验，为你挑选了3个好方法。 ... [详细]
程序员
Laravel - 从设置表中设置全局变量

如何解决《Laravel-从设置表中设置全局变量》经验，为你挑选了1个好方法。 ... [详细]
程序员
检查给定变量的最佳方法是NaN与否？

如何解决《检查给定变量的最佳方法是NaN与否？》经验，为你挑选了1个好方法。 ... [详细]
程序员
Java 8 CompletableFuture与Netty Future

如何解决《Java8CompletableFuture与NettyFuture》经验，为你挑选了0个好方法。 ... [详细]
程序员
可以将以下IF-ELSE块压缩成单个IF语句吗？

如何解决《可以将以下IF-ELSE块压缩成单个IF语句吗？》经验，为你挑选了2个好方法。 ... [详细]
程序员
在多对多关系中的多个连接条件

如何解决《在多对多关系中的多个连接条件》经验，为你挑选了0个好方法。 ... [详细]
程序员
如何将数据表值从整数转换为字符串

如何解决《如何将数据表值从整数转换为字符串》经验，为你挑选了1个好方法。 ... [详细]
程序员
Beego控制器中的JSON响应

如何解决《Beego控制器中的JSON响应》经验，为你挑选了1个好方法。 ... [详细]

周扒pi

这个屌丝很懒，什么也没留下！

关注作者

Tags | 热门标签

RankList | 热门文章