# Which parser generator are you using (if any)?

**URL:** https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233
**Category:** Parsing
**Created:** [February 12, 2020, 11:58am UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233 "2020-02-12T11:58:49Z")
**Posts on this page:** 20
**Page:** 2

<div class="post-metadata">

### Author: ![jurgen.vinju](https://d.strumenta.community/user_avatar/d.strumenta.community/jurgen.vinju/32/20_2.png) [@jurgen.vinju](https://d.strumenta.community/u/jurgen.vinju)
#### Post date: [February 15, 2020, 10:27am UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/21 "2020-02-15T10:27:41Z")

</div>

Ah right. Well with glr and gll algorithms the parser can accept any context-free grammar rules. So the grammar does not have to be in LR form or LL form.

This means that the BNF formalism can have modularity features, composable grammars, because you can just throw rules together and the parser will still “work”.

This also means you can write: `E ::= E “+” E | E “*” E | ... ; ` etc; lets say for C you have 20 to 30 rules, without additional nonterminals to get the grammar into an LR or LL shape.

So grammars for glr and gll have a lot fewer nonterminals and they can be composed and reused without change.

There is a cost however for this; first the parser time is non optimal (but usually still nearly linear for programming languages). Second a grammar might be ambiguous and so you get more than one tree perhaps. The “BNF” caters for this with additional constraints such as “\>“ between rules for operator precedence and whitespace constraints for defining the offside rule and such. Nevertheless ambiguity and the fact that you have to find out about it is the main cost to pay for the simplicity and modularity of general context free grammars.

---

<div class="post-metadata">

### Author: ![jurgen.vinju](https://d.strumenta.community/user_avatar/d.strumenta.community/jurgen.vinju/32/20_2.png) [@jurgen.vinju](https://d.strumenta.community/u/jurgen.vinju)
#### Post date: [February 15, 2020, 10:29am UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/22 "2020-02-15T10:29:37Z")

</div>

Thanks for the link to the paper btw!

---

<div class="post-metadata">

### Author: ![anon67755252](https://d.strumenta.community/letter_avatar_proxy/v4/letter/a/e9c0ed/32.png) [@anon67755252](https://d.strumenta.community/u/anon67755252)
#### Post date: [February 15, 2020, 11:15am UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/23 "2020-02-15T11:15:30Z")

</div>

I’m doing the same thing within an LR grammar:

Exp : Primary  
| Exp ‘+’ Exp \*\> add\_   
| Exp ‘-’ Exp _\> sub\_   
| Exp '_’ Exp \*\> mul\_   
| Exp ‘/’ Exp \*\> div\_   
;  
It’s the disambiguating rules that make this work properly.  
LRstar is free and comtains 9 sample grammars:  
[http://lrstar.tech](http://lrstar.tech)

---

<div class="post-metadata">

### Author: ![jurgen.vinju](https://d.strumenta.community/user_avatar/d.strumenta.community/jurgen.vinju/32/20_2.png) [@jurgen.vinju](https://d.strumenta.community/u/jurgen.vinju)
#### Post date: [February 15, 2020, 11:38am UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/24 "2020-02-15T11:38:35Z")

</div>

Nice! That’s what we want indeed. I’ll read about it in the paper right? Composing grammars too?

---

<div class="post-metadata">

### Author: ![jurgen.vinju](https://d.strumenta.community/user_avatar/d.strumenta.community/jurgen.vinju/32/20_2.png) [@jurgen.vinju](https://d.strumenta.community/u/jurgen.vinju)
#### Post date: [February 15, 2020, 11:40am UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/25 "2020-02-15T11:40:39Z")

</div>

I guess your transformations not only disambiguate but also make the grammar deterministic. Right?

---

<div class="post-metadata">

### Author: ![anon67755252](https://d.strumenta.community/letter_avatar_proxy/v4/letter/a/e9c0ed/32.png) [@anon67755252](https://d.strumenta.community/u/anon67755252)
#### Post date: [February 15, 2020, 7:34pm UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/26 "2020-02-15T19:34:40Z")

</div>

Yes, the disambiguating rules make the grammar LALR(1) or LR(1)  
and therefore deterministic. The disambiguating rules are simple:

/\* Operator precedence. \*/

{ ‘==’ ‘!=’ } \<\< // Lowest priority.  
{ ‘+’ ‘-’ } \<\<  
{ ‘\*’ ‘/’ } \<\< // Highest priority.

The parsing speed is very fast, reading a 227,00 line file in 0.095 seconds  
and it builds a symbol table and AST. And the parsers handles the “typedef”  
problem automatically. And there is more …  
.

---

<div class="post-metadata">

### Author: ![jurgen.vinju](https://d.strumenta.community/user_avatar/d.strumenta.community/jurgen.vinju/32/20_2.png) [@jurgen.vinju](https://d.strumenta.community/u/jurgen.vinju)
#### Post date: [February 15, 2020, 7:57pm UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/27 "2020-02-15T19:57:33Z")

</div>

Nice to meet another parsing expert and enthusiast!

---

<div class="post-metadata">

### Author: ![anon67755252](https://d.strumenta.community/letter_avatar_proxy/v4/letter/a/e9c0ed/32.png) [@anon67755252](https://d.strumenta.community/u/anon67755252)
#### Post date: [February 15, 2020, 8:32pm UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/28 "2020-02-15T20:32:45Z")

</div>

Thanks. I’ve been in the shadows for 30 years, an unknown, unappreciated, etc.

---

<div class="post-metadata">

### Author: ![jurgen.vinju](https://d.strumenta.community/user_avatar/d.strumenta.community/jurgen.vinju/32/20_2.png) [@jurgen.vinju](https://d.strumenta.community/u/jurgen.vinju)
#### Post date: [February 15, 2020, 8:46pm UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/29 "2020-02-15T20:46:58Z")

</div>

It happens to more parsing people. If you are ever near Amsterdam, pls feel free to visit!

---

<div class="post-metadata">

### Author: ![anon67755252](https://d.strumenta.community/letter_avatar_proxy/v4/letter/a/e9c0ed/32.png) [@anon67755252](https://d.strumenta.community/u/anon67755252)
#### Post date: [February 15, 2020, 9:06pm UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/30 "2020-02-15T21:06:07Z")

</div>

I have Dutch ancestors, but I’ve never been to the Netherlands. I’m half European,  
but stuck in the US for now. The linguistic part of computer science is a very neglected  
area, just like the usability and readability areas.

---

<div class="post-metadata">

### Author: ![cristian.vasile](https://d.strumenta.community/letter_avatar_proxy/v4/letter/c/958977/32.png) [@cristian.vasile](https://d.strumenta.community/u/cristian.vasile)
#### Post date: [February 16, 2020, 7:16am UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/31 "2020-02-16T07:16:16Z")

</div>

> [@anon67755252](#):
>
> I have created a parser generator

Paul,  
Could you create a **commercial** tool able to read a grammar file and emits valid but random source code? Some sort of [CSmith](https://embed.cs.utah.edu/csmith/) but not bind to C language. A generic tool.  
Such a tool should be very useful for stress testing DSL languages or anything which can have a grammar (protocols, smart contracts etc)

Also I would like to point you to an interesting match between data visualization and semantics:  
[Large scale semantic representation with flame graphs](https://www.anlp.jp/proceedings/annual_meeting/2015/pdf_dir/C1-5.pdf) Institute for Excellence in Higher Education, Tohoku University.  
The authors of the paper are using a type of vizualisation called [flame graph](http://www.brendangregg.com/flamegraphs.html) invented by Brendan Gregg.

---

<div class="post-metadata">

### Author: ![anon67755252](https://d.strumenta.community/letter_avatar_proxy/v4/letter/a/e9c0ed/32.png) [@anon67755252](https://d.strumenta.community/u/anon67755252)
#### Post date: [February 16, 2020, 7:56am UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/32 "2020-02-16T07:56:46Z")

</div>

Yes, I think the generated parser could have a built-in random symbol creator which  
selects the next valid symbol from the valid symbols of the current state. Then go to  
the designated state and repeat until EOF is chosen. It would need some way to make  
it stop, because the number of choices could be huge.

---

<div class="post-metadata">

### Author: ![cristian.vasile](https://d.strumenta.community/letter_avatar_proxy/v4/letter/c/958977/32.png) [@cristian.vasile](https://d.strumenta.community/u/cristian.vasile)
#### Post date: [February 16, 2020, 8:42pm UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/33 "2020-02-16T20:42:11Z")

</div>

Paul,

I would like to suggest other difficult area you can explore:  
A code migration tool able to transpile from language A to language B  
C -\> GO  
FORTRAN -\> Julia  
Oracle PL/SQL -\> Java & SQL  
Oracle SQL -\> Microsoft SQL  
GO -\> Rust  
C -\> Rust ex: see [C to Rust tool / C2Rust](https://immunant.com/blog/2019/08/introduction-to-c2rust/)

---

<div class="post-metadata">

### Author: ![anon67755252](https://d.strumenta.community/letter_avatar_proxy/v4/letter/a/e9c0ed/32.png) [@anon67755252](https://d.strumenta.community/u/anon67755252)
#### Post date: [February 16, 2020, 9:31pm UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/34 "2020-02-16T21:31:49Z")

</div>

Yes, that is one of my goals, to translate from one language to another.  
I have the tool now, (LRstar, [http://lrstar.tech](http://lrstar.tech)) . It only took 30 years  
to perfect it, without (much) pay.  
Now, I need to see some money (venture capitalists or angel investors),  
before I spend another hour on this stuff.

---

<div class="post-metadata">

### Author: ![ftomassetti](https://d.strumenta.community/user_avatar/d.strumenta.community/ftomassetti/32/8_2.png) [@ftomassetti](https://d.strumenta.community/u/ftomassetti)
#### Post date: [February 17, 2020, 6:58am UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/35 "2020-02-17T06:58:13Z")

</div>

> [@cristian.vasile](#):
>
> Oracle PL/SQL → Java & SQL

I have seen that there is some interest on this one. We have also made an article on that one: [Convert PL/SQL code to Java](https://tomassetti.me/convert-pl-sql-code-to-java/)

---

<div class="post-metadata">

### Author: ![igor.dejanovic](https://d.strumenta.community/user_avatar/d.strumenta.community/igor.dejanovic/32/43_2.png) [@igor.dejanovic](https://d.strumenta.community/u/igor.dejanovic)
#### Post date: [February 17, 2020, 10:40am UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/36 "2020-02-17T10:40:39Z")

</div>

I couldn’t agree more. BNF like languages are the languages of our domain. Not only it is easier to implement and maintain parsers using these DSLs but is also easier to communicate ideas, improve consistency, capture the domain knowledge etc., usual benefits of applying DSLs. In my opinion it is well worth a time investment in learning these DSLs and related tools if you are planing to work in this field. Manual parser implementation is interesting, in my view, if you want to learn more in-depth of how parsers work.

---

<div class="post-metadata">

### Author: ![cristian.vasile](https://d.strumenta.community/letter_avatar_proxy/v4/letter/c/958977/32.png) [@cristian.vasile](https://d.strumenta.community/u/cristian.vasile)
#### Post date: [February 17, 2020, 4:44pm UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/37 "2020-02-17T16:44:43Z")

</div>

I remember that article, it is an interview.

---

<div class="post-metadata">

### Author: ![cristian.vasile](https://d.strumenta.community/letter_avatar_proxy/v4/letter/c/958977/32.png) [@cristian.vasile](https://d.strumenta.community/u/cristian.vasile)
#### Post date: [February 17, 2020, 5:15pm UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/38 "2020-02-17T17:15:50Z")

</div>

Paul,

Just some ideas in random order able to pursuit without too much cash.  
o setup a “cloud” transpiler service, for low cost high volume of source files based on number of LOCs

o your tool is a perfect match for databases, parsing SQL dialects is notorious hard. There is a wave of so called GPU databases like [Omnisci](https://www.omnisci.com/), [Kinetica](https://www.kinetica.com/), [SQream](https://sqream.com/) etc. Omnsici was a bison/yacc shop and switched to Apache Calcite despite the fact their kernel is written in C++; SQream is using an in house solution written in Haskell; Kinetica - i do not know but you can approach them with a business proposal. Last time when I checked they were at sql 92 something ?!?!

o HPC - Cray inc, had developed a language called [Chapel](https://chapel-lang.org/) dedicated for parallel programming; could be an option to translate FORTRAN libraries to Chapel under a professional service agreement.

o promote your tool(s) in a broader way

o [RISC V](https://www.sifive.com/) processors; There is a new open source processor called RISC V/five and you can help their customers on testing process.

European initiatives on HPC field.  
o [EuroHPC](http://eurohpc.eu/) is a joint collaboration between European countries and the European Union about developing and supporting exascale supercomputing by 2022/2023

o [European Processor Initiative](https://sipearl.com/)

Hope this helps.

---

<div class="post-metadata">

### Author: ![alessio.stalla](https://d.strumenta.community/user_avatar/d.strumenta.community/alessio.stalla/32/361_2.png) [@alessio.stalla](https://d.strumenta.community/u/alessio.stalla)
#### Post date: [February 17, 2020, 5:47pm UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/39 "2020-02-17T17:47:28Z")

</div>

I believe the reverse (Java+SQL -\> stored procedures) could also be commercially interesting. Stored procedures are not a bad idea per se, but languages for storage procedures are quite bad.

---

<div class="post-metadata">

### Author: ![anon67755252](https://d.strumenta.community/letter_avatar_proxy/v4/letter/a/e9c0ed/32.png) [@anon67755252](https://d.strumenta.community/u/anon67755252)
#### Post date: [February 17, 2020, 8:45pm UTC](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233/40 "2020-02-17T20:45:47Z")

</div>

Thanks for the suggestions.

[Previous page](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233.md?page=1)

[Next page](https://d.strumenta.community/t/which-parser-generator-are-you-using-if-any/233.md?page=3)
