Python NLTK: Bigrams trigrams fourgrams

python nltk n-gram

If you apply some set theory (if I'm interpreting your question correctly), you'll see that the trigrams you want are simply elements [2:5], [4:7], [6:8], etc. of the token list.

You could generate them like this:

>>> new_trigrams = []>>> c = 2>>> while c < len(token) - 2:...     new_trigrams.append((token[c], token[c+1], token[c+2]))...     c += 2>>> print new_trigrams[('are', 'you', '?'), ('?', 'i', 'am'), ('am', 'fine', 'and')]

python nltk n-gram

Try everygrams:

from nltk import everygramslist(everygrams('hello', 1, 5))

[out]:

[('h',), ('e',), ('l',), ('l',), ('o',), ('h', 'e'), ('e', 'l'), ('l', 'l'), ('l', 'o'), ('h', 'e', 'l'), ('e', 'l', 'l'), ('l', 'l', 'o'), ('h', 'e', 'l', 'l'), ('e', 'l', 'l', 'o'), ('h', 'e', 'l', 'l', 'o')]

Word tokens:

from nltk import everygramslist(everygrams('hello word is a fun program'.split(), 1, 5))

[out]:

[('hello',), ('word',), ('is',), ('a',), ('fun',), ('program',), ('hello', 'word'), ('word', 'is'), ('is', 'a'), ('a', 'fun'), ('fun', 'program'), ('hello', 'word', 'is'), ('word', 'is', 'a'), ('is', 'a', 'fun'), ('a', 'fun', 'program'), ('hello', 'word', 'is', 'a'), ('word', 'is', 'a', 'fun'), ('is', 'a', 'fun', 'program'), ('hello', 'word', 'is', 'a', 'fun'), ('word', 'is', 'a', 'fun', 'program')]

python nltk n-gram

I do it like this:

def words_to_ngrams(words, n, sep=" "):    return [sep.join(words[i:i+n]) for i in range(len(words)-n+1)]

This takes a list of words as input and returns a list of ngrams (for given n), separated by sep (in this case a space).

CodeHunter

Python NLTK: Bigrams trigrams fourgrams

Recent Posts

How can I color dots in a xy scatterplot according to column value?

How to update a claim in ASP.NET Identity?

What does {0} mean when initializing an object?

Accessing members of items in a JSONArray with Java

How to log SQL statements in Spring Boot?

Powershell Get-WebSite name parameter is ignored

How to detect scroll to bottom of html element

Java synchronized method

How to test controllers with CodeIgniter?

Detect Visual Composer

Matplotlib: Specify format of floats for tick labels

Rails join a list of strings with commas and "and" before the last