public class DutchWordTokenizer extends WordTokenizer
Constructor and Description |
---|
DutchWordTokenizer() |
Modifier and Type | Method and Description |
---|---|
String |
getTokenizingCharacters() |
List<String> |
tokenize(String text)
Tokenizes just like WordTokenizer with the exception for words such as
"oma's" that contain an apostrophe in their middle.
|
getProtocols, isEMail, isUrl, joinEMails, joinEMailsAndUrls, joinUrls
public List<String> tokenize(String text)
tokenize
in interface Tokenizer
tokenize
in class WordTokenizer
text
- Text to tokenizepublic String getTokenizingCharacters()
getTokenizingCharacters
in class WordTokenizer