## Abstract

The Dyck language, which consists of well-balanced sequences of parentheses, is one of the most fundamental context-free languages. The Dyck edit distance quantifies the number of edits (character insertions, deletions, and substitutions) required to make a given length-n parenthesis sequence well-balanced. RNA Folding involves a similar problem, where a closing parenthesis can match an opening parenthesis of the same type irrespective of their ordering. For example, in RNA Folding, both () and)(are valid matches, whereas the Dyck language only allows () as a match. Both of these problems have been studied extensively in the literature. Using fast matrix multiplication, it is possible to compute their exact solutions in time O(n^{2.687}) (Chi, Duan, Xie, Zhang, STOC'22), and a (1 + ϵ)-multiplicative approximation is known with a running time of Ω(n^{2.372}). The impracticality of fast matrix multiplication often makes combinatorial algorithms much more desirable. Unfortunately, it is known that the problems of (exactly) computing the Dyck edit distance and the folding distance are at least as hard as Boolean matrix multiplication. Thereby, they are unlikely to admit truly subcubic-time combinatorial algorithms. In terms of fast approximation algorithms that are combinatorial in nature, the state of the art for Dyck edit distance is an O(log n)-factor approximation algorithm that runs in near-linear time (Saha, FOCS'14), whereas for RNA Folding only an ϵn-additive approximation in (Equation presented) time (Saha, FOCS'17) is known. In this paper, we make substantial improvements to the state of the art for Dyck edit distance (with any number of parenthesis types). We design a constant-factor approximation algorithm that runs in Õ(n^{1.971}) time (the first constant-factor approximation in subquadratic time). Moreover, we develop a (1 + ϵ)-factor approximation algorithm running in (Equation presented) time, which improves upon the earlier additive approximation. Finally, we design a (3 + ϵ)-approximation that takes (Equation presented) time, where d ≥ 1 is an upper bound on the sought distance. As for RNA folding, for any s ≥ 1, we design a factor-s approximation algorithm that runs in O(n + (n/s)^{3}) time. To the best of our knowledge, this is the first nontrivial approximation algorithm for RNA Folding that can go below the n^{2} barrier. All our algorithms are combinatorial in nature.

Original language | English (US) |
---|---|

Title of host publication | 49th EATCS International Conference on Automata, Languages, and Programming, ICALP 2022 |

Editors | Mikolaj Bojanczyk, Emanuela Merelli, David P. Woodruff |

Publisher | Schloss Dagstuhl- Leibniz-Zentrum fur Informatik GmbH, Dagstuhl Publishing |

ISBN (Electronic) | 9783959772358 |

DOIs | |

State | Published - Jul 1 2022 |

Event | 49th EATCS International Conference on Automata, Languages, and Programming, ICALP 2022 - Paris, France Duration: Jul 4 2022 → Jul 8 2022 |

### Publication series

Name | Leibniz International Proceedings in Informatics, LIPIcs |
---|---|

Volume | 229 |

ISSN (Print) | 1868-8969 |

### Conference

Conference | 49th EATCS International Conference on Automata, Languages, and Programming, ICALP 2022 |
---|---|

Country/Territory | France |

City | Paris |

Period | 7/4/22 → 7/8/22 |

## All Science Journal Classification (ASJC) codes

- Software