---
title: "Introduction to tidyEmoji"
author: 'Youzhi Yu<br><span style="font-weight: 400; font-size: 0.85em;">University of Chicago</span>'
output:
  rmarkdown::html_vignette:
    toc: true
    toc_depth: 3
vignette: >
  %\VignetteIndexEntry{Introduction to tidyEmoji}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  message = FALSE,
  warning = FALSE,
  fig.width = 8,
  fig.height = 5,
  comment = "#>"
)
```

## Overview

Emoji are everywhere in modern text (social-media posts, product reviews, chat
and support logs, survey free-text), and they carry information that plain words
do not. Yet summarising emoji from a corpus is surprisingly awkward. Unicode does
not interact cleanly with regular expressions, not every code point is an emoji,
and a single visible emoji is often built from several code points joined
together. Counting "how many posts contain an emoji" or "which emoji are most
common" by hand quickly becomes painful.

**tidyEmoji** removes that friction. It provides a small family of verbs that
take a data frame and the name of a text column, and return tidy data frames
that drop straight into a `dplyr`/`ggplot2` workflow:

| Task | Function(s) |
|------|-------------|
| Summarise / filter | `emoji_summary()`, `emoji_filter()` |
| Extract | `emoji_extract_nest()`, `emoji_extract_unnest()`, `emoji_tokens()` |
| Count | `emoji_frequency()`, `top_n_emojis()` |
| Categorise | `emoji_categorize()` |
| Score sentiment | `emoji_sentiment()` |
| Score emotions | `emoji_emotion()`, `emoji_emotion_label()` |
| Custom lexicons | `emoji_lexicons()`, `register_emoji_lexicon()`, `emoji_score()` |
| Translate | `emoji_to_text()`, `text_to_emoji()`, `as_emoji*()` |
| Search | `emoji_search()` |
| Relate | `emoji_pairs()`, `emoji_cooccurrence()`, `emoji_ngrams()` |
| Measure | `emoji_position()`, `emoji_density()`, `emoji_ratio()` |
| Interpretation risk | `emoji_ambiguity()`, `emoji_risk()`, `emoji_flag_ambiguous()` |
| Context | `emoji_context()`, `emoji_collocations()` |
| Functional type | `emoji_type()`, `emoji_faceness()`, `as_emoji_type()` |
| Time | `emoji_trend()`, `emoji_turnover()`, `emoji_seasonality()`, `emoji_version_profile()`, `emoji_adoption_lag()` |
| Text-emoji mismatch | `emoji_incongruity()`, `emoji_congruence()`, `emoji_incongruity_profile()` |
| Model features | `emoji_dfm()` |
| LLM pipelines | `emoji_sanitize()`, `emoji_token_cost()` |
| Provenance | `emoji_provenance()`, `emoji_unicode_version()`, `emoji_unicode_releases()` |

Two design choices are worth highlighting:

* **Grapheme-aware detection.** Detection is performed on whole grapheme
  clusters, so skin-tone modifiers (👍🏽) and zero-width-joiner sequences such
  as the family emoji (👨‍👩‍👧‍👦) are treated as a *single* emoji rather than
  being split into their component parts. This is illustrated in the
  [extraction section](#a-note-on-grapheme-aware-detection).
* **Tidy by default.** Every verb returns a tibble, follows the
  `verb(data, text_column)` convention, and supports unquoted column names, so
  the functions compose naturally with the pipe. Grouping composes too: the
  verbs that work a row at a time carry a `group_by()` through to their result,
  as `dplyr::mutate()` and `dplyr::filter()` do, while the ones that pool
  across rows -- the counts, the co-occurrence tables, the time series -- warn
  that they ignore it and return one corpus-wide answer. See `?tidyEmoji` for
  the full contract.

```{r setup}
library(tidyEmoji)
library(dplyr)

# The charts below need three packages that tidyEmoji only Suggests. A vignette
# has to build without its optional dependencies, so the plotting chunks are
# gated on this flag rather than assuming the packages are there.
has_plot_pkgs <- all(vapply(
  c("ggplot2", "forcats", "stringr"),
  requireNamespace, logical(1), quietly = TRUE
))
if (has_plot_pkgs) library(ggplot2)
```

That gate covers building this vignette. It does *not* carry over to the
script `knitr` tangles out of it, the one
`edit(vignette("introduction", package = "tidyEmoji"))` opens: a chunk's
`eval` option is a build-time instruction and is not part of the extracted
code, so `introduction.R` calls `ggplot()` unconditionally. Everything in it
that uses tidyEmoji alone runs against the package's declared dependencies;
install \pkg{ggplot2}, \pkg{forcats} and \pkg{stringr} if you want to run
the whole file.

## Example data

Throughout this vignette we use a sample of text collected in Atlanta,
Georgia. The data happens to come from a social-media corpus, but nothing below
is specific to any platform: any data frame with a text column will do.

```{r}
# read.csv() rather than readr::read_csv(): readr is a Suggests package, and
# reading one CSV does not need it. system.file() rather than a bare path, so
# this chunk works from the installed introduction.R as well as at build time
# -- the corpus lives in inst/extdata and is installed with the package.
ata_tweets <- tibble::as_tibble(
  utils::read.csv(
    system.file("extdata", "ata_tweets.csv", package = "tidyEmoji"),
    encoding = "UTF-8", stringsAsFactors = FALSE
  )
)
ata_tweets
```

The actual text lives in the `full_text` column, which is the column we pass to
each tidyEmoji verb.

## Detecting and summarising emoji

### `emoji_summary()`

`emoji_summary()` answers the first question one usually asks of a new corpus:
*how much emoji is in here?* It returns a one-row tibble with the number of
entries that contain at least one emoji and the total number of entries. An
entry is counted once regardless of how many emoji it holds.

```{r}
summary_tbl <- ata_tweets %>%
  emoji_summary(full_text)

summary_tbl
```

Here, `r format(summary_tbl$n_with_emoji, big.mark = ",")` of the
`r format(summary_tbl$n_total, big.mark = ",")` entries
(`r round(100 * summary_tbl$n_with_emoji / summary_tbl$n_total, 1)`%)
contain at least one emoji.

### `emoji_filter()`

`emoji_filter()` keeps only the rows whose text contains at least one emoji,
preserving every original column. This is useful when you want to compare
emoji-bearing and emoji-free text, or restrict an analysis to the emoji subset.
(`emoji_tweets()` is a synonym retained for backward compatibility.)

```{r}
ata_tweets %>%
  emoji_filter(full_text)
```

## Extracting emoji

tidyEmoji offers three complementary ways to pull the emoji out of text,
depending on the shape of output you want.

### `emoji_extract_nest()`

`emoji_extract_nest()` leaves the data unchanged except for an added
list-column, `.emoji_unicode`, holding the emoji found in each row. The original
data structure is preserved, which makes this convenient as an intermediate
step.

```{r}
ata_tweets %>%
  emoji_extract_nest(full_text) %>%
  select(full_text, .emoji_unicode)
```

### `emoji_extract_unnest()`

`emoji_extract_unnest()` returns a long, tidy table with one row per
(entry, emoji) pair: `.row_number` records the position of the entry in the data,
`.emoji_unicode` is the emoji, and `.emoji_count` is how many times that emoji
occurs in that entry. Entries without emoji are dropped.

```{r}
emoji_per_tweet <- ata_tweets %>%
  emoji_extract_unnest(full_text)

emoji_per_tweet
```

We can use this to plot how many emoji each emoji-bearing entry contains:

```{r, eval = has_plot_pkgs, fig.alt = "Bar chart of the number of emoji per emoji-bearing entry. About two-thirds of entries contain a single emoji, with a long, thin tail of more emoji-heavy entries."}
emoji_per_tweet %>%
  group_by(.row_number) %>%
  summarise(n_emoji = sum(.emoji_count)) %>%
  ggplot(aes(n_emoji)) +
  geom_bar() +
  scale_x_continuous(breaks = seq(1, 15)) +
  labs(x = "Number of emoji in the entry",
       y = "Number of entries",
       title = "Most emoji-bearing entries contain a single emoji")
```

About two-thirds of emoji-bearing entries carry just one emoji, with a
long, thin tail of more emoji-heavy entries.

### `emoji_tokens()`

`emoji_tokens()` produces a "one row per emoji occurrence" table, the emoji
analogue of a tidy-text token table. It keeps the original columns and adds the
glyph (`.emoji`) together with its name (`.emoji_name`), category
(`.emoji_category`) and sentiment score (`.emoji_sentiment`). This single call
gives you everything needed for counting, joining and plotting.

```{r}
ata_tweets %>%
  emoji_tokens(full_text)
```

### A note on grapheme-aware detection

Modern emoji are frequently composed of several code points: a base emoji plus a
skin-tone modifier, or several emoji joined by zero-width joiners. tidyEmoji
detects emoji at the level of grapheme clusters, so these stay intact. The
example below contains exactly two emoji (one family and one thumbs-up), and
tidyEmoji counts them as such rather than splitting the family into four people
or separating the thumb from its skin tone:

```{r}
demo <- data.frame(
  text = c("our family \U0001F468‍\U0001F469‍\U0001F467‍\U0001F466",
           "great work \U0001F44D\U0001F3FD")
)

demo %>%
  emoji_extract_unnest(text)
```

## Counting emoji across the corpus

### `emoji_frequency()`

`emoji_frequency()` counts how often each emoji appears across the whole text
column (an entry containing the same emoji twice contributes 2) and returns the
result sorted by descending count, annotated with each emoji's name, shortcode
and category.

```{r}
ata_tweets %>%
  emoji_frequency(full_text)
```

### `top_n_emojis()`

When you only need the leaders, `top_n_emojis()` is a convenience wrapper around
`emoji_frequency()` that returns the `n` most frequent emoji (default `n = 20`).

```{r}
top_20_emojis <- ata_tweets %>%
  top_n_emojis(full_text)

top_20_emojis
```

Plotting the top 20, coloured by category, gives an immediate sense of how the
community expresses itself:

```{r, eval = has_plot_pkgs, fig.alt = "Horizontal bar chart of the 20 most frequent emoji in the corpus, coloured by Unicode category."}
top_20_emojis %>%
  mutate(emoji_name = stringr::str_replace_all(emoji_name, "_", " "),
         emoji_name = forcats::fct_reorder(emoji_name, n)) %>%
  ggplot(aes(n, emoji_name, fill = emoji_category)) +
  geom_col() +
  labs(x = "Count",
       y = NULL,
       fill = "Category",
       title = "The 20 most frequent emoji")
```

The `unicode` column holds the actual glyph, should you wish to render the emoji
themselves on a plot (this requires a graphics device with an emoji-capable
font). You can also request a different number of emoji:

```{r, eval = has_plot_pkgs, fig.alt = "Horizontal bar chart of the 10 most frequent emoji in the corpus, coloured by Unicode category."}
ata_tweets %>%
  top_n_emojis(full_text, n = 10) %>%
  mutate(emoji_name = stringr::str_replace_all(emoji_name, "_", " "),
         emoji_name = forcats::fct_reorder(emoji_name, n)) %>%
  ggplot(aes(n, emoji_name, fill = emoji_category)) +
  geom_col() +
  labs(x = "Count", y = NULL, fill = "Category",
       title = "The 10 most frequent emoji")
```

## Categorising emoji

The Unicode standard organises emoji into 10 categories (see
`?category_unicode_crosswalk`). `emoji_categorize()` keeps the emoji-bearing rows
and adds a `.emoji_category` column listing the distinct categories present in
each row, separated by `|` when a row spans more than one.

```{r}
ata_emoji_category <- ata_tweets %>%
  emoji_categorize(full_text) %>%
  select(.emoji_category)

ata_emoji_category
```

We can tally the most common category combinations:

```{r, eval = has_plot_pkgs, fig.alt = "Horizontal bar chart of the most common emoji category combinations that appear in more than 20 entries."}
ata_emoji_category %>%
  count(.emoji_category, sort = TRUE) %>%
  filter(n > 20) %>%
  mutate(.emoji_category = forcats::fct_reorder(.emoji_category, n)) %>%
  ggplot(aes(n, .emoji_category)) +
  geom_col() +
  labs(x = "Number of entries", y = NULL,
       title = "Most common emoji category combinations")
```

To count the 10 individual categories rather than their combinations, split the
`.emoji_category` strings on `|` with `tidyr::separate_longer_delim()`:

```{r, eval = has_plot_pkgs, fig.alt = "Horizontal bar chart of how often each individual Unicode emoji category is used, dominated by Smileys & Emotion followed by People & Body."}
ata_emoji_category %>%
  tidyr::separate_longer_delim(.emoji_category, delim = "|") %>%
  count(.emoji_category, sort = TRUE) %>%
  mutate(.emoji_category = forcats::fct_reorder(.emoji_category, n)) %>%
  ggplot(aes(n, .emoji_category)) +
  geom_col() +
  labs(x = "Number of entries", y = NULL,
       title = "Emoji category usage")
```

"Smileys & Emotion" dominates, followed by "People & Body". Note that an entry
spanning several categories is counted once in each, so these counts can exceed
the number of emoji-bearing entries.

### `emoji_type()`: faces versus everything else

The consumer-behaviour literature repeatedly contrasts *emotional* (face) emoji
with *semantic* (object) emoji and finds that the two have different effects on
engagement. `emoji_type()` recodes the Unicode group and subgroup into that
smaller functional vocabulary, and `emoji_faceness()` reduces it to a single
per-entry share:

```{r}
ata_tweets %>%
  emoji_type(full_text) %>%
  count(.emoji_type, sort = TRUE) %>%
  head(5)

ata_tweets %>%
  emoji_faceness(full_text) %>%
  summarise(mean_faceness = mean(.emoji_faceness, na.rm = TRUE),
            all_faces = sum(.emoji_faceness == 1, na.rm = TRUE))
```

`as_emoji_type()` is the vector-level version, for typing a glyph you already
have in hand.

## Scoring emoji sentiment

### `emoji_sentiment()`

Emoji are a strong sentiment signal, and `emoji_sentiment()` surfaces it
directly. It adds `.emoji_n` (the number of emoji in the entry),
`.emoji_n_scored` (the number that appear in the lexicon), and
`.emoji_sentiment` (the mean sentiment of the scored emoji, from -1 to +1).
Scores come from the bundled `emoji_sentiment_lexicon` (described below);
entries with no emoji, or whose emoji are not in the lexicon, receive `NA`.

```{r}
ata_sentiment <- ata_tweets %>%
  emoji_sentiment(full_text)

ata_sentiment %>%
  select(.emoji_n, .emoji_sentiment)
```

### Sentiment distribution

Looking across the entries that contain at least one scored emoji:

```{r, eval = has_plot_pkgs, fig.alt = "Histogram of the mean emoji sentiment per entry, which is concentrated on the positive side of the scale."}
ata_sentiment %>%
  filter(!is.na(.emoji_sentiment)) %>%
  ggplot(aes(.emoji_sentiment)) +
  geom_histogram(binwidth = 0.1) +
  labs(x = "Mean emoji sentiment",
       y = "Number of entries",
       title = "Emoji sentiment skews positive")
```

As is typical of social-media text, emoji sentiment leans strongly positive.

### Sentiment by category

Because `emoji_tokens()` attaches a sentiment score to every emoji occurrence, we
can summarise average sentiment by category in a couple of lines:

```{r, eval = has_plot_pkgs, fig.alt = "Horizontal bar chart of the average emoji sentiment within each Unicode category."}
ata_tweets %>%
  emoji_tokens(full_text) %>%
  group_by(.emoji_category) %>%
  summarise(mean_sentiment = mean(.emoji_sentiment, na.rm = TRUE),
            n_scored = sum(!is.na(.emoji_sentiment))) %>%
  filter(n_scored > 0) %>%
  mutate(.emoji_category = forcats::fct_reorder(.emoji_category, mean_sentiment)) %>%
  ggplot(aes(mean_sentiment, .emoji_category)) +
  geom_col() +
  labs(x = "Mean sentiment", y = NULL,
       title = "Average emoji sentiment by category")
```

### The sentiment lexicon

The scores come from `emoji_sentiment_lexicon`, the *Emoji Sentiment Ranking* of
Kralj Novak et al. (2015), computed from around 70,000 tweets annotated in 13
European languages. You can work with it directly, for instance to find the
most positive and most negative reasonably common emoji:

```{r}
emoji_sentiment_lexicon %>%
  filter(occurrences >= 500) %>%
  slice_max(sentiment_score, n = 8) %>%
  select(emoji, unicode_name, occurrences, sentiment_score)

emoji_sentiment_lexicon %>%
  filter(occurrences >= 500) %>%
  slice_min(sentiment_score, n = 8) %>%
  select(emoji, unicode_name, occurrences, sentiment_score)
```

## Interpretation risk: how much do readers disagree?

The most practically important fact about emoji is that people do not agree on
what they mean. Miller et al. (2016) showed the *same* rendering to many
readers and found they disagreed about whether it was positive, neutral or
negative around a quarter of the time.

That disagreement was already inside the package. The Emoji Sentiment Ranking
keeps the raw `negative`/`neutral`/`positive` annotation counts behind its
collapsed score, which is an empirical interpretation distribution per glyph.
`emoji_ambiguity()` reads it out:

```{r}
emoji_ambiguity() %>%
  filter(n_annotations > 500) %>%
  head(5)
```

Because the measure is a property of the glyph, the corpus-level question --
"which of *my* emoji are most likely to be misread?" -- is one call:

```{r}
ata_tweets %>%
  emoji_flag_ambiguous(full_text, top_n = 5)
```

`emoji_risk()` is the per-entry version, and the pair `.emoji_n` /
`.emoji_n_scored` keeps the coverage honest: an entry whose emoji are absent
from the lexicon is not a low-risk entry, it is an unmeasured one.

```{r}
ata_tweets %>%
  emoji_risk(full_text) %>%
  filter(.emoji_n > 1) %>%
  select(.emoji_n, .emoji_n_scored, .emoji_ambiguity_mean,
         .emoji_n_ambiguous) %>%
  head(5)
```

The same counts also put an error bar on the sentiment score itself. A glyph
annotated eight times should not carry the authority of one annotated eight
thousand times, and `emoji_sentiment(se = TRUE)` says so:

```{r}
ata_tweets %>%
  emoji_sentiment(full_text, se = TRUE) %>%
  filter(!is.na(.emoji_sentiment)) %>%
  select(.emoji_n_scored, .emoji_sentiment, .emoji_sentiment_se) %>%
  head(5)
```

## Scoring emoji emotions

Valence (negative↔positive) is only one affective dimension. `emoji_emotion()`
goes further, scoring each entry's emoji across the eight Plutchik emotions
(anger, anticipation, disgust, fear, joy, sadness, surprise, trust) using the
bundled EmoTag1200 lexicon (Shoeb & de Melo, 2020). Scores are in `[0, 1]`.

```{r}
ata_emotion <- ata_tweets %>%
  emoji_emotion(full_text)

ata_emotion %>%
  select(.emoji_joy, .emoji_trust, .emoji_anger, .emoji_n)
```

A quick way to read the result is the dominant emotion per entry:

```{r}
ata_tweets %>%
  emoji_emotion_label(full_text) %>%
  count(.emoji_emotion, sort = TRUE)
```

The emotion scores join through the same codepoint-normalised key as sentiment,
so emoji carrying the `U+FE0F` variation selector resolve correctly.

## Bringing your own lexicon

Sentiment and emotion scoring share one pluggable engine. `emoji_lexicons()`
lists the bundled lexicons (plus any you have registered):

```{r}
emoji_lexicons()
```

`emoji_score()` is the generic scorer underneath `emoji_sentiment()`: give it
any data frame with an emoji column and a score column (say, scores tailored
to your own domain), and it returns the per-row mean, joined through the same
codepoint-normalised key as everything else:

```{r}
my_lexicon <- data.frame(
  emoji = c("\U0001f600", "\U0001f621", "\U0001f637"),
  score = c(1, -1, -0.5)
)

data.frame(text = c("great \U0001f600", "bad \U0001f621\U0001f637", "none")) %>%
  emoji_score(text, lexicon = my_lexicon)
```

`register_emoji_lexicon()` stores a lexicon under a name for the session, so
you can refer to it in `emoji_score()`, or in `emoji_emotion()` if it carries
emotion columns:

```{r}
register_emoji_lexicon("mine", my_lexicon)
emoji_lexicons() %>% filter(name == "mine")
```

## Relating emoji: co-occurrence and sequences

Which emoji appear *together*? `emoji_pairs()` returns a tidy edge list: one
row per pair of distinct emoji that co-occur in the same entry, with the number
of entries in which they do. The `item1`/`item2`/`n` shape matches
`widyr::pairwise_count()` and feeds directly into graph tools such as igraph,
tidygraph and ggraph:

```{r}
emoji_edges <- ata_tweets %>%
  emoji_pairs(full_text)

emoji_edges
```

The strongest pairings make a readable chart on their own:

```{r, eval = has_plot_pkgs, fig.alt = "Horizontal bar chart of the most frequent emoji pairs, labelled by the two glyphs of each pair."}
emoji_edges %>%
  slice_max(n, n = 10) %>%
  mutate(pair = paste(item1, item2),
         pair = forcats::fct_reorder(pair, n)) %>%
  ggplot(aes(n, pair)) +
  geom_col() +
  labs(x = "Number of entries containing both", y = NULL,
       title = "Emoji that appear together")
```

Set `directed = TRUE` to order each pair by first appearance, or supply
`doc_id` to pool several rows (a conversation, a user, a day) into one
document. `emoji_cooccurrence(diagonal = TRUE)` additionally returns each
emoji's document frequency on the diagonal.

Order also matters *within* an entry. `emoji_ngrams()` slides a window over
each entry's emoji in reading order (any text in between is ignored), which is
the raw material for sequence and Markov-style analyses:

```{r}
ata_tweets %>%
  emoji_ngrams(full_text) %>%
  count(.emoji_ngram, sort = TRUE)
```

All the relational verbs canonicalise glyphs through the same
codepoint-normalised key as the rest of the package, so qualified and
unqualified forms of one emoji count as a single node.

## The words around an emoji

Which emoji occurred is only half the story: emoji are polysemous, and their
reading is decided by the co-text. `emoji_context()` returns one row per
occurrence with a window of the surrounding text. Other emoji are blanked out
of the window, so a neighbouring glyph never leaks into it:

```{r}
ata_tweets %>%
  emoji_context(full_text, window = 4) %>%
  select(.row_number, .emoji, .emoji_context) %>%
  head(5)
```

Aggregated over a corpus, those windows give the emoji-word associations that
a sense inventory would otherwise have to supply -- derived from your own data,
so they are neither stale nor licence-encumbered:

```{r}
ata_tweets %>%
  emoji_collocations(full_text, window = 4, min_n = 5) %>%
  head(10)
```

Tokenisation stops at whitespace on purpose. If you need stemming or stopword
removal, hand the result to `tidytext` rather than expecting this verb to grow
a tokeniser.

## Measuring how emoji are used

*Where* emoji sit and *how much* of the text they occupy are studied signals in
their own right. `emoji_position()` reports each entry's first and last emoji
position and the mean relative position from 0 (start) to 1 (end):

```{r, eval = has_plot_pkgs, fig.alt = "Histogram of the mean relative position of emoji within each entry, showing emoji concentrated towards the end of the text."}
ata_tweets %>%
  emoji_position(full_text) %>%
  filter(!is.na(.emoji_rel_position)) %>%
  ggplot(aes(.emoji_rel_position)) +
  geom_histogram(binwidth = 0.05) +
  labs(x = "Mean relative position of the entry's emoji",
       y = "Number of entries",
       title = "Emoji cluster at the end of a message")
```

`emoji_density()` normalises the emoji count by text length (per character and
per whitespace-delimited token), and `emoji_ratio()` reports what share of the
text's characters belong to emoji, including an `.emoji_only` flag for entries
that are nothing but emoji (and whitespace):

```{r}
ata_tweets %>%
  emoji_ratio(full_text) %>%
  summarise(
    n_emoji_only = sum(.emoji_only, na.rm = TRUE),
    mean_ratio   = mean(.emoji_ratio[.emoji_ratio > 0], na.rm = TRUE)
  )
```

## Emoji over time

Almost every substantive emoji study is longitudinal. Our sample has no
timestamp, so the dates below are *synthetic* -- they illustrate the mechanics,
not a finding:

```{r}
dated <- ata_tweets %>%
  mutate(posted_at = as.Date("2021-01-01") + (seq_len(n()) - 1) %% 540)

dated %>%
  emoji_trend(full_text, posted_at, by = "quarter", top_n = 3)
```

`emoji_turnover()` asks a different question -- not how often an emoji is used
but how much of the *repertoire* changes from one period to the next:

```{r}
dated %>%
  emoji_turnover(full_text, posted_at, by = "quarter")
```

Two of the time verbs read their time axis off the reference table, which
already records the Unicode version that introduced each glyph.
`emoji_version_profile()` needs no timestamp at all as a result, so "how new is
this corpus's emoji vocabulary?" is a single call:

```{r}
ata_tweets %>%
  emoji_version_profile(full_text) %>%
  head(8)
```

With real timestamps, `emoji_adoption_lag()` goes one step further and compares
each glyph's first appearance in the corpus with its Unicode release date --
see `emoji_unicode_releases()` for that lookup, and `emoji_seasonality()` for
month-of-year, day-of-week and hour-of-day cycles.

## Emoji as model features

For classification and regression work, `emoji_dfm()` turns the corpus into a
document-by-emoji feature table: one row per entry (or per `doc_id`), one
column per emoji, weighted by counts, binary presence or tf-idf. Every entry is
kept (emoji-free rows are all zeros), so the table binds row-for-row to your
outcome columns:

```{r}
ata_tweets %>%
  emoji_dfm(full_text, weighting = "tfidf") %>%
  select(1:6)
```

## Text-emoji mismatch

Two separate literatures converge on one statistic. NLP sarcasm detection uses
emoji-text sentiment incongruity as a feature; marketing research finds that a
mismatch between a review's words and its emoji lowers perceived helpfulness
and authenticity. Both want the same number.

tidyEmoji deliberately does not score text -- that choice belongs in your
script, where the method is visible. Here is a deliberately crude word-list
scorer standing in for `tidytext` + AFINN, `sentimentr` or a transformer:

```{r}
positive <- c("love", "great", "best", "happy", "good", "thanks", "beautiful")
negative <- c("hate", "worst", "bad", "sad", "awful", "sick", "tired")

scored <- ata_tweets %>%
  mutate(text_score = vapply(
    strsplit(tolower(full_text), "[^a-z]+"),
    function(w) as.numeric(sum(w %in% positive) - sum(w %in% negative)),
    numeric(1)
  ))
```

`scale` has no default: AFINN runs -5 to 5, VADER -1 to 1, and a model's logits
on nothing in particular, so you have to say how the two sides were made
comparable. `"rank"` puts both on percentiles and is the safest choice.

The percentiles are taken over the rows the comparison is defined on -- those
with both a scorable emoji and a text score -- not over the whole corpus. Only
373 of these 2000 tweets qualify, and ranking the text score against all 2000
would compare a percentile of one population with a percentile of another. One
consequence is a useful sanity check: on the rank scale the mean gap over the
scored rows is exactly zero, so a non-zero mean means your filtering, not your
data.

```{r}
incong <- scored %>%
  emoji_incongruity(full_text, text_score, scale = "rank")

incong %>%
  filter(!is.na(.emoji_incongruity)) %>%
  count(.emoji_polarity_flip)

incong %>%
  filter(.emoji_polarity_flip) %>%
  select(full_text, .emoji_sentiment, text_score) %>%
  head(3)
```

Note what happens to entries with no scorable emoji: they get `NA`, never `0`.
A neutral emoji and no emoji at all are different states, and collapsing them
biases every model downstream. `emoji_incongruity_profile()` aggregates the
same numbers by glyph, and `emoji_congruence()` is the identical engine under
the marketing framing.

```{r}
scored %>%
  emoji_incongruity_profile(full_text, text_score, scale = "rank", min_n = 10)
```

## Translating emoji to and from text

Replacing emoji with words is useful for accessibility (screen readers) and as
an NLP normalisation step before tokenising. `emoji_to_text()` does this for a
whole column, in either Unicode-name or shortcode form; `text_to_emoji()` is the
inverse.

```{r}
demo <- data.frame(text = "great \U0001f600 love \u2764\ufe0f")
demo %>% emoji_to_text(text, format = "name")
demo %>% emoji_to_text(text, format = "shortcode")
demo %>%
  emoji_to_text(text, format = "shortcode") %>%
  text_to_emoji(text)
```

Note that the qualified heart (which carries the `U+FE0F` variation selector)
translates just as reliably as any other emoji, thanks to the normalised join
key. For ad-hoc, vector-level use there are also `as_emoji_name()`,
`as_emoji_shortcode()` and `as_emoji()`:

```{r}
as_emoji_name(c("\U0001f600", "\u2764\ufe0f"))
as_emoji_shortcode(c("\U0001f600", "\u2764\ufe0f"))
as_emoji(c("grinning", "heart"))
```

## Searching the emoji catalogue

`emoji_search()` looks emoji up by keyword, name or shortcode
(case-insensitive, literal matching), returning a tidy tibble you can filter
further or feed into the other verbs:

```{r}
emoji_search("happy")
emoji_search("celebration")
```

## Emoji in language-model pipelines

Emoji are now a preprocessing decision in every LLM pipeline, and the decision
matters: they inflate token counts several-fold, and models disambiguate them
poorly. `emoji_token_cost()` gives the exact sizes plus a clearly-labelled
estimate -- pass your real tokeniser through `tokenizer` when the number
matters:

```{r}
ata_tweets %>%
  emoji_token_cost(full_text) %>%
  filter(.emoji_n > 0) %>%
  summarise(emoji = sum(.emoji_n),
            bytes = sum(.emoji_bytes),
            codepoints = sum(.emoji_codepoints),
            est_tokens = sum(.emoji_token_estimate))
```

`emoji_sanitize()` then applies one *named* policy to the column. The
capability is not new -- most of it exists across `emoji_to_text()` and the
extraction verbs -- but the named argument shows up in a script diff and in a
methods section, which is the point:

```{r}
demo_llm <- data.frame(text = "ship it \U0001f680 today")
for (p in c("keep", "strip", "name", "shortcode", "placeholder")) {
  cat(format(p, width = 12), emoji_sanitize(demo_llm, text, policy = p)$text,
      "\n")
}
```

## Recording provenance

"Emoji" is not a fixed object. Which glyphs exist, what they are called and
which lexicon scored them all depend on versions, and a result is not
reproducible without them. `emoji_provenance()` puts the lot in one row you can
paste into a methods section:

```{r}
emoji_provenance() %>% glimpse()
```

## Bundled datasets

tidyEmoji ships four datasets, each documented with its own help page:

* **`emoji_sentiment_lexicon`**: emoji sentiment scores from the Emoji
  Sentiment Ranking (see `?emoji_sentiment_lexicon`).
* **`emoji_emotion_lexicon`**: emoji emotion scores from EmoTag1200
  (see `?emoji_emotion_lexicon`).
* **`emoji_unicode_crosswalk`**: one row per (name, glyph) pair, mapping
  names / shortcodes to glyphs and categories. The mapping is many-to-many
  both ways, so join on `key` rather than `emoji_name` unless you want the
  duplicates (see `?emoji_unicode_crosswalk`).
* **`category_unicode_crosswalk`**: one row per Unicode category, listing its
  emoji.

These are regenerated from the current Unicode emoji list by the scripts in the
package's `data-raw/` directory.

## References

Kralj Novak P, Smailović J, Sluban B, Mozetič I (2015). Sentiment of Emojis.
*PLoS ONE* 10(12): e0144296.
<https://doi.org/10.1371/journal.pone.0144296>. The Emoji Sentiment Ranking is
distributed under the Creative Commons Attribution-ShareAlike 4.0 International
(CC BY-SA 4.0) licence.

Miller H, Thebault-Spieker J, Chang S, Johnson I, Terveen L, Hecht B (2016).
"Blissfully Happy" or "Ready to Fight": Varying Interpretations of Emoji.
*ICWSM 2016*. The source of the disagreement result behind
`emoji_ambiguity()`.

Shoeb AAM, de Melo G (2020). EmoTag1200: Understanding the Association between
Emojis and Emotions. *EMNLP 2020*.
<https://aclanthology.org/2020.emnlp-main.720/>. The EmoTag1200 data is
distributed under the MIT licence.
